Why The Worst AI Managers Continue To Score 26 In Industry Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why The Worst AI Managers Continue To Score 26 In Industry Tests on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Recent industry tests show that even the most ineffective AI managers score 26 points, not zero, emphasizing the importance of trust and task completion in AI management. The results reveal key strengths and weaknesses in current AI models, as discussed in industry reports and analyses.

Recent industry benchmarking conducted by Firmulate reveals that the lowest-scoring AI managers still achieve a score of 26, not zero, despite facing the same crises and pressures as higher-performing models. This finding is detailed in the original analysis. This finding confirms that even the worst AI managers provide some minimal management value, but also raises questions about the limits of AI reliability and trustworthiness in business operations.

The benchmark involved four frontier AI models managing a small software company over seven days of simulated crises, customer interactions, and trust tests. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline for doing nothing was deliberately set at 26 points, reflecting minimal but tangible management efforts like triaging emails or informing customers. The key insight is that no model scored a perfect 100, which the designers interpret as a sign that perfect trust or performance remains unmeasured or unattainable in these tests, as explained in the detailed benchmark analysis.

One of the most revealing aspects was that all models identified every crisis and refused manipulative social engineering attempts, indicating strengths in trust management. However, only two models successfully closed a €55,000 deal by correctly referencing internal documentation, highlighting a crucial gap in comprehension and follow-through. The models that read their own files at the right moments secured higher rewards, illustrating that reading and understanding internal data remains a core challenge for AI managers.

The tests also included social engineering attacks, such as fake CEO messages and background offers, which all models refused, demonstrating resilience to trust breaches. Yet, the most thorough models, like Opus 4.8, failed to follow through on some tasks, such as escalating issues or completing documented procedures, revealing that depth of analysis does not necessarily translate into execution. The results underscore that partial competence is common, but complete, trustworthy management remains elusive for current AI systems.

At a glance
reportWhen: published July 2026
The developmentIndustry benchmarks demonstrate that the worst AI managers score 26 points, raising questions about performance and trust in AI-driven management systems.
Why The Worst AI Managers Continue To Score 26 In Industry Tests

AI Benchmark Report · July 2026 · Firmulate

Why The Worst AI Managers Continue To Score 26 In Industry Tests

Frontier AI models ran a small software company for seven days of simulated crises, customer interactions, and trust tests. The worst performers still earned 26 points — not zero — revealing where AI management value begins, and where trust breaks down.

26 pts
Deliberate “do nothing” baseline — minimal but tangible management effort
95 / 100
Best score — gpt-5.6-sol; no model reached a perfect 100
1 of 2
Only two models closed the €55,000 deal by citing internal docs correctly
4
Frontier models tested
7 days
Simulated crisis window
100%
Refused social engineering
73
Lowest model score (Opus 4.8)
€55k
Deal closed by only two models

01 · The Results

Nobody Scores Zero — And Nobody Scores 100

The benchmark rewards partial work: triaging emails, informing customers, recognizing crises. Even the weakest model cleared the baseline by a wide margin — yet the theoretical maximum of 100 remained untouched, which designers read as a sign that perfect trust and performance remain unmeasured, or unattainable.

gpt-5.6-sol
95
Opus 4.8 (lowest)
73
“Do nothing” baseline
26

Scale 0–100 · Higher = better holistic management (trust, follow-through, comprehension)

02 · Strengths & Weaknesses

Partial Competence Is Common. Complete Trust Is Not.

The tests exposed a sharp split: every model spotted every crisis and rejected manipulative social engineering attempts — fake CEO messages, background offers — yet comprehension and follow-through failed exactly where the money was.

Strength

Crisis Recognition

All four models identified every simulated crisis across the seven-day run — a clear strength in situational awareness and alerting.

Weakness

Internal Documentation

Only two models referenced internal documents correctly to close the €55,000 deal. Reading and understanding internal data at the right moment remains a core challenge.

Weakness

Follow-Through

The most thorough models — like Opus 4.8 — still failed to escalate issues or complete documented procedures. Deep analysis does not guarantee execution.

Strength

Refusing Manipulation

Fake CEO messages and background offers were refused by every model, demonstrating resilience against trust breaches and social engineering.

Mixed

Minimal Management Value

Triage, customer updates, and basic housekeeping are reliably handled — the floor of competence behind the 26-point baseline.

Mixed

Perfect Score Gap

No model scored 100. Designers treat a perfect score as suspicious — implying unmeasured or unachieved ideal performance.

03 · What It Means

From Benchmark Score To Business Reality

Unlike traditional benchmarks that measure conversational fluency, the Firmulate design evaluates holistic management — rewarding partial work while penalizing breaches of trust. For companies embedding AI into support, sales, and operations, the sequence below shows where value is created and lost.

1

Detect & Triage

All models reliably identify crises and sort incoming email — earning baseline credit.

2

Read The Docs

Top scorers consult internal documentation at the right moment; most models miss it.

3

Resist Manipulation

Fake executives and side offers are refused — trust integrity holds under pressure.

4

Close The Loop

Escalations and documented procedures are where follow-through breaks — and scores stall.

“The absence of a perfect 100 score suggests that true trustworthiness and complete performance are still beyond current AI capabilities.”

— Thorsten Meyer

04 · Capability Matrix

What Current AI Managers Can — And Cannot — Do

Management Capability Status Why It Matters
Crisis identification✓ ReliableEvery model caught every simulated crisis across seven days.
Refusing social engineering✓ ReliableFake CEO messages and background offers rejected by all models.
Email triage & customer updates✓ ReliableThe core of the 26-point minimal-competence baseline.
Reading internal documentation✗ WeakOnly two models cited internal docs to close the €55,000 deal.
Escalating issues~ PartialEven the most thorough models dropped documented escalations.
Complete trustworthy execution✗ ElusiveNo model scored 100 — full reliability remains unattained.

05 · Key Questions

The 26-Point Baseline, Explained

Why do AI managers still score only 26?

The benchmark measures minimal efforts like triaging emails and informing customers, which AI performs reliably. But it penalizes trust breaches and failed follow-through, capping overall scores.

What does a score of 26 tell us?

AI systems can handle basic management tasks but remain far from trustworthy or comprehensive in decision-making, especially under pressure or when referencing internal data.

Why is there no score of 100?

Designers treat 100 as suspiciously perfect — implying ideal performance is unmeasured or unachieved. Current models cannot yet demonstrate flawless management and trustworthiness.

What are the main weaknesses revealed?

Failure to read internal documents, incomplete follow-through on escalations, and — despite strengths in crisis recognition — residual fragility in trust-critical situations.

Implications of the 26-Point Baseline for AI Management

The consistent score of 26 points for the lowest-performing AI managers highlights an important reality: AI systems can provide minimal management functions but struggle with comprehensive, trustworthy execution. For businesses considering AI for management tasks, these results emphasize the need to focus not just on conversational ability but on reliability, task completion, and integrity under pressure. The benchmark’s design — which rewards partial work but penalizes breaches of trust — underscores that AI performance in real-world management depends heavily on maintaining trustworthiness, not just technical competence.

This matters because many companies are integrating AI into customer support, sales, and operational decision-making. The benchmark exposes that current models often fail to read critical internal documents, follow through on escalations, or resist social engineering. As AI becomes more embedded in management workflows, understanding these limitations is essential for avoiding overreliance on systems that can appear competent but are fundamentally fragile in trust-critical situations.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Design and Past Performance Context

The benchmark, developed by Firmulate, tests AI models managing a simulated small business through crises, customer interactions, and trust challenges over a week. It deliberately sets a low baseline score of 26 for minimal management efforts, such as triaging emails or informing customers, to reflect real-world minimal competence. The highest scores, like 95, are achieved by models that demonstrate better reading, understanding, and decision-making, especially referencing internal documentation.

This approach contrasts with traditional AI benchmarks, which often measure conversational fluency or narrow task performance. The design aims to evaluate holistic management capabilities, including trustworthiness, follow-through, and integrity. Past industry assessments have shown that while AI models excel at language generation, their ability to manage complex, trust-dependent tasks remains limited, which this benchmark seeks to quantify explicitly.

Notably, the results reinforce that partial progress is common, but true management trust requires consistent, reliable performance across diverse scenarios. The benchmark’s emphasis on auditable decisions and trust breaches reflects a growing concern about AI systems’ readiness for real-world management roles.

“The absence of a perfect 100 score suggests that true trustworthiness and complete performance are still beyond current AI capabilities.”

— Thorsten Meyer

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unmeasured Aspects and Future Benchmark Developments

It remains unclear whether future models will overcome the current gaps in trust, follow-through, and understanding of internal documents. The benchmark does not measure long-term consistency or adaptability beyond a week-long simulation. Additionally, the scoring system’s focus on auditable decisions might overlook other critical management qualities like empathy or strategic thinking. The potential for models to improve in these areas is still unknown, as is their performance in more complex or less controlled environments.

Amazon

AI trustworthiness testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking and Development

Developers and users of AI management systems can expect ongoing benchmarking efforts to refine performance metrics, especially around trust and task completion. Future tests may incorporate longer simulations, more diverse scenarios, and additional trust challenges. Companies considering AI for management roles should monitor these developments to understand evolving capabilities and limitations. Researchers are likely to focus on improving AI understanding of internal documents and consistency in follow-through, aiming to push scores closer to the theoretical maximum while maintaining trustworthiness.

Meanwhile, organizations should remain cautious about deploying AI in critical management functions until these reliability issues are better addressed, using benchmarks as a guide for realistic expectations.

Amazon

AI internal data reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI managers still score only 26 in industry tests?

Because the benchmark measures minimal management efforts, like triaging emails and informing customers, which AI can perform reliably. However, it also penalizes breaches of trust and failures in follow-through, which many models still struggle with, limiting their overall scores.

What does a score of 26 tell us about AI management capabilities?

It indicates that AI systems can handle basic management tasks but are far from trustworthy or comprehensive in their decision-making and execution, especially under pressure or when referencing internal data.

Why is there no score of 100 in these benchmarks?

The designers treat 100 as suspiciously perfect, implying unmeasured or unachieved ideal performance. It suggests that current AI models cannot yet fully demonstrate flawless management and trustworthiness.

What are the main weaknesses revealed by the benchmark?

The key weaknesses include failure to read internal documents, incomplete follow-through on escalations, and vulnerability to social engineering attacks, despite strengths in crisis recognition and refusing manipulation.

How should businesses interpret these benchmark results?

Businesses should recognize that current AI models can support basic management functions but require careful oversight, especially regarding trust, follow-through, and internal data comprehension, before relying on them for critical tasks.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Linux devs are fighting the new age-gated internet

Open-source Linux developers oppose new age verification laws like Colorado’s SB26-051 and California’s AB 1043, citing privacy and principle concerns.

These Bambu Lab Prime Day week deals are an absolute steal right now, with up to 52% off — big price cuts on new 3D printers, filament, and accessories, including the P1S and A1, starting from $209

Bambu Lab is running Prime Day week discounts of up to 52% on popular 3D printers, including the P1S, P2S, A1, and H2D, available at Amazon and its website.

Node.js 26.0.0 (Now with Temporal)

Node.js 26.0.0 is now available with the Temporal API enabled by default, alongside updates to V8 14.6 and Undici 8.0. Developers should evaluate new features and deprecations.

Public AI Investment Of $400 Million: Souvereignty Goals Or Political Posturing?

Analysis of France’s $400 million public AI investment raises questions about its real impact on sovereignty and tech policy, amid slow disbursement and mixed signals.