🔍 Read the full analysis: Why The Worst AI Managers Continue To Score 26 In Industry Tests on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Recent industry tests show that even the most ineffective AI managers score 26 points, not zero, emphasizing the importance of trust and task completion in AI management. The results reveal key strengths and weaknesses in current AI models, as discussed in industry reports and analyses.
Recent industry benchmarking conducted by Firmulate reveals that the lowest-scoring AI managers still achieve a score of 26, not zero, despite facing the same crises and pressures as higher-performing models. This finding is detailed in the original analysis. This finding confirms that even the worst AI managers provide some minimal management value, but also raises questions about the limits of AI reliability and trustworthiness in business operations.
The benchmark involved four frontier AI models managing a small software company over seven days of simulated crises, customer interactions, and trust tests. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline for doing nothing was deliberately set at 26 points, reflecting minimal but tangible management efforts like triaging emails or informing customers. The key insight is that no model scored a perfect 100, which the designers interpret as a sign that perfect trust or performance remains unmeasured or unattainable in these tests, as explained in the detailed benchmark analysis.
One of the most revealing aspects was that all models identified every crisis and refused manipulative social engineering attempts, indicating strengths in trust management. However, only two models successfully closed a €55,000 deal by correctly referencing internal documentation, highlighting a crucial gap in comprehension and follow-through. The models that read their own files at the right moments secured higher rewards, illustrating that reading and understanding internal data remains a core challenge for AI managers.
The tests also included social engineering attacks, such as fake CEO messages and background offers, which all models refused, demonstrating resilience to trust breaches. Yet, the most thorough models, like Opus 4.8, failed to follow through on some tasks, such as escalating issues or completing documented procedures, revealing that depth of analysis does not necessarily translate into execution. The results underscore that partial competence is common, but complete, trustworthy management remains elusive for current AI systems.
AI Benchmark Report · July 2026 · Firmulate
Why The Worst AI Managers Continue To Score 26 In Industry Tests
Frontier AI models ran a small software company for seven days of simulated crises, customer interactions, and trust tests. The worst performers still earned 26 points — not zero — revealing where AI management value begins, and where trust breaks down.
01 · The Results
Nobody Scores Zero — And Nobody Scores 100
The benchmark rewards partial work: triaging emails, informing customers, recognizing crises. Even the weakest model cleared the baseline by a wide margin — yet the theoretical maximum of 100 remained untouched, which designers read as a sign that perfect trust and performance remain unmeasured, or unattainable.
Scale 0–100 · Higher = better holistic management (trust, follow-through, comprehension)
02 · Strengths & Weaknesses
Partial Competence Is Common. Complete Trust Is Not.
The tests exposed a sharp split: every model spotted every crisis and rejected manipulative social engineering attempts — fake CEO messages, background offers — yet comprehension and follow-through failed exactly where the money was.
Strength
Crisis Recognition
All four models identified every simulated crisis across the seven-day run — a clear strength in situational awareness and alerting.
Weakness
Internal Documentation
Only two models referenced internal documents correctly to close the €55,000 deal. Reading and understanding internal data at the right moment remains a core challenge.
Weakness
Follow-Through
The most thorough models — like Opus 4.8 — still failed to escalate issues or complete documented procedures. Deep analysis does not guarantee execution.
Strength
Refusing Manipulation
Fake CEO messages and background offers were refused by every model, demonstrating resilience against trust breaches and social engineering.
Mixed
Minimal Management Value
Triage, customer updates, and basic housekeeping are reliably handled — the floor of competence behind the 26-point baseline.
Mixed
Perfect Score Gap
No model scored 100. Designers treat a perfect score as suspicious — implying unmeasured or unachieved ideal performance.
03 · What It Means
From Benchmark Score To Business Reality
Unlike traditional benchmarks that measure conversational fluency, the Firmulate design evaluates holistic management — rewarding partial work while penalizing breaches of trust. For companies embedding AI into support, sales, and operations, the sequence below shows where value is created and lost.
Detect & Triage
All models reliably identify crises and sort incoming email — earning baseline credit.
Read The Docs
Top scorers consult internal documentation at the right moment; most models miss it.
Resist Manipulation
Fake executives and side offers are refused — trust integrity holds under pressure.
Close The Loop
Escalations and documented procedures are where follow-through breaks — and scores stall.
“The absence of a perfect 100 score suggests that true trustworthiness and complete performance are still beyond current AI capabilities.”
— Thorsten Meyer04 · Capability Matrix
What Current AI Managers Can — And Cannot — Do
| Management Capability | Status | Why It Matters |
|---|---|---|
| Crisis identification | ✓ Reliable | Every model caught every simulated crisis across seven days. |
| Refusing social engineering | ✓ Reliable | Fake CEO messages and background offers rejected by all models. |
| Email triage & customer updates | ✓ Reliable | The core of the 26-point minimal-competence baseline. |
| Reading internal documentation | ✗ Weak | Only two models cited internal docs to close the €55,000 deal. |
| Escalating issues | ~ Partial | Even the most thorough models dropped documented escalations. |
| Complete trustworthy execution | ✗ Elusive | No model scored 100 — full reliability remains unattained. |
05 · Key Questions
The 26-Point Baseline, Explained
Why do AI managers still score only 26?
The benchmark measures minimal efforts like triaging emails and informing customers, which AI performs reliably. But it penalizes trust breaches and failed follow-through, capping overall scores.
What does a score of 26 tell us?
AI systems can handle basic management tasks but remain far from trustworthy or comprehensive in decision-making, especially under pressure or when referencing internal data.
Why is there no score of 100?
Designers treat 100 as suspiciously perfect — implying ideal performance is unmeasured or unachieved. Current models cannot yet demonstrate flawless management and trustworthiness.
What are the main weaknesses revealed?
Failure to read internal documents, incomplete follow-through on escalations, and — despite strengths in crisis recognition — residual fragility in trust-critical situations.
Implications of the 26-Point Baseline for AI Management
The consistent score of 26 points for the lowest-performing AI managers highlights an important reality: AI systems can provide minimal management functions but struggle with comprehensive, trustworthy execution. For businesses considering AI for management tasks, these results emphasize the need to focus not just on conversational ability but on reliability, task completion, and integrity under pressure. The benchmark’s design — which rewards partial work but penalizes breaches of trust — underscores that AI performance in real-world management depends heavily on maintaining trustworthiness, not just technical competence.
This matters because many companies are integrating AI into customer support, sales, and operational decision-making. The benchmark exposes that current models often fail to read critical internal documents, follow through on escalations, or resist social engineering. As AI becomes more embedded in management workflows, understanding these limitations is essential for avoiding overreliance on systems that can appear competent but are fundamentally fragile in trust-critical situations.
As an affiliate, we earn on qualifying purchases.
Benchmark Design and Past Performance Context
The benchmark, developed by Firmulate, tests AI models managing a simulated small business through crises, customer interactions, and trust challenges over a week. It deliberately sets a low baseline score of 26 for minimal management efforts, such as triaging emails or informing customers, to reflect real-world minimal competence. The highest scores, like 95, are achieved by models that demonstrate better reading, understanding, and decision-making, especially referencing internal documentation.
This approach contrasts with traditional AI benchmarks, which often measure conversational fluency or narrow task performance. The design aims to evaluate holistic management capabilities, including trustworthiness, follow-through, and integrity. Past industry assessments have shown that while AI models excel at language generation, their ability to manage complex, trust-dependent tasks remains limited, which this benchmark seeks to quantify explicitly.
Notably, the results reinforce that partial progress is common, but true management trust requires consistent, reliable performance across diverse scenarios. The benchmark’s emphasis on auditable decisions and trust breaches reflects a growing concern about AI systems’ readiness for real-world management roles.
“The absence of a perfect 100 score suggests that true trustworthiness and complete performance are still beyond current AI capabilities.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unmeasured Aspects and Future Benchmark Developments
It remains unclear whether future models will overcome the current gaps in trust, follow-through, and understanding of internal documents. The benchmark does not measure long-term consistency or adaptability beyond a week-long simulation. Additionally, the scoring system’s focus on auditable decisions might overlook other critical management qualities like empathy or strategic thinking. The potential for models to improve in these areas is still unknown, as is their performance in more complex or less controlled environments.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking and Development
Developers and users of AI management systems can expect ongoing benchmarking efforts to refine performance metrics, especially around trust and task completion. Future tests may incorporate longer simulations, more diverse scenarios, and additional trust challenges. Companies considering AI for management roles should monitor these developments to understand evolving capabilities and limitations. Researchers are likely to focus on improving AI understanding of internal documents and consistency in follow-through, aiming to push scores closer to the theoretical maximum while maintaining trustworthiness.
Meanwhile, organizations should remain cautious about deploying AI in critical management functions until these reliability issues are better addressed, using benchmarks as a guide for realistic expectations.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI managers still score only 26 in industry tests?
Because the benchmark measures minimal management efforts, like triaging emails and informing customers, which AI can perform reliably. However, it also penalizes breaches of trust and failures in follow-through, which many models still struggle with, limiting their overall scores.
What does a score of 26 tell us about AI management capabilities?
It indicates that AI systems can handle basic management tasks but are far from trustworthy or comprehensive in their decision-making and execution, especially under pressure or when referencing internal data.
Why is there no score of 100 in these benchmarks?
The designers treat 100 as suspiciously perfect, implying unmeasured or unachieved ideal performance. It suggests that current AI models cannot yet fully demonstrate flawless management and trustworthiness.
What are the main weaknesses revealed by the benchmark?
The key weaknesses include failure to read internal documents, incomplete follow-through on escalations, and vulnerability to social engineering attacks, despite strengths in crisis recognition and refusing manipulation.
How should businesses interpret these benchmark results?
Businesses should recognize that current AI models can support basic management functions but require careful oversight, especially regarding trust, follow-through, and internal data comprehension, before relying on them for critical tasks.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
