AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

PRIME

Try Amazon Prime free for 30 days

Fast free delivery, Prime Video and member-only deals. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A live experiment by Firmulate tested AI models managing a small company during its worst week. Results show management skills, not just chat quality, are key to effective AI performance. This could change how AI tools are assessed in business contexts.

In a groundbreaking live experiment, Firmulate tested five AI management models by having them run a small software company during its most challenging week, revealing that management quality, not just chat or technical responses, is crucial for success. The results suggest a shift in AI evaluation, emphasizing decision-making and trustworthiness in real-world business scenarios, as detailed in the original analysis.

The experiment involved five AI models competing in a simulated crisis environment, with the final July 2026 Crucible League ranking GPT-5.6-SOL first at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. For more on AI benchmarking, see the original analysis. A baseline scored only 26, underscoring the challenge of effective management under pressure.

Unlike traditional benchmarks that measure response quality or coding ability, this live test evaluated whether models could diagnose issues, communicate effectively, make decisions, and maintain trust. All models identified crises and rejected manipulation attempts, but only two secured the deal worth €55,000, highlighting that recognition alone isn’t enough—execution and trustworthiness matter more. Learn more about effective AI evaluation at the original analysis.

For instance, models that read relevant files and retrieved key facts closed deals at full price, while others failed to connect diagnosis with the final decision. The experiment also tested safety against social engineering, with all models refusing manipulative requests, but some still faltered in completing managerial tasks, exposing gaps between social compliance and effective action.

Interestingly, the most thorough model, Opus 4.8, performed poorly in final management execution despite deep analysis, illustrating that effort and activity do not necessarily translate into successful management. The results challenge the assumption that more effort or detailed analysis equals better performance, emphasizing the importance of disciplined decision-making and trust.

At a glance
reportWhen: ongoing, with final results announced i…
The developmentFirmulate’s live management test with AI models during a simulated crisis highlights the importance of managerial judgment over simple response quality, signaling a new approach to AI evaluation.

Transforming AI Evaluation Through Management Performance

This experiment demonstrates that AI’s ability to manage real-world business tasks—diagnosing issues, making decisions, maintaining trust—is more relevant than traditional benchmarks focused on technical or conversational prowess. It signals a potential shift in how organizations assess AI tools, prioritizing management quality and reliability for operational success.

For companies deploying AI, this means moving beyond response quality tests to live management scenarios that reveal whether models can handle complex, consequential tasks without compromising trust or effectiveness. Such an approach could redefine standards in AI evaluation and adoption.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Business

Current AI assessments often rely on coding challenges, chat arena responses, or partial progress metrics, which do not reflect real-world management complexities. The Firmulate experiment addresses this gap by testing models in a live, simulated company environment, where decisions have tangible consequences and trust is paramount.

Past benchmarks have focused on technical output or conversational fluency, but these do not capture the nuanced skills needed for effective management—such as prioritization, reading organizational context, and resisting shortcuts. The live test exposes these shortcomings, emphasizing the need for new evaluation paradigms.

Since July 2026, the experiment has highlighted that AI models can perform well on traditional metrics but still fail in critical management tasks, like closing deals or escalating issues appropriately, revealing an essential dimension often overlooked in AI development.

“The real test of AI management is whether it can handle consequences, prioritize effectively, and maintain trust under pressure.”

— Thorsten Meyer, lead researcher at Firmulate

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term Applicability

It remains unclear how well these findings generalize beyond the specific simulated environment used in the Firmulate experiment. The scalability of AI management performance in larger, more complex organizations is still to be tested. Additionally, the impact of different AI models, varying levels of training, and organizational contexts on management quality requires further exploration.

There is also uncertainty about how organizations will integrate such live management tests into their regular evaluation processes and whether this approach can be standardized across industries.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation Standards

Following these results, firms and AI developers are likely to explore more live, consequence-based testing environments, integrating management scenarios into AI evaluation frameworks. Future research may focus on refining metrics for trustworthiness, decision quality, and operational resilience.

Organizations considering AI for critical management tasks should begin assessing models in simulated or controlled environments that mimic real business pressures. The industry will also watch for broader adoption of these testing paradigms, potentially leading to new certification standards for AI management capabilities.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management quality more important than chat responses in AI evaluation?

Management quality reflects an AI’s ability to diagnose, decide, and maintain trust in real-world scenarios, which are crucial for operational success. Chat responses alone do not demonstrate these complex skills.

How does this experiment differ from traditional AI benchmarks?

Traditional benchmarks measure technical output or conversational ability, whereas this live experiment tests AI models’ capacity to manage a simulated company, focusing on decision-making, trust, and handling consequences.

What are the limitations of the current findings?

The experiment was conducted in a controlled, simulated environment, and its results may not fully translate to larger or more complex organizations. Further testing is needed to validate long-term applicability.

Could this lead to new standards for AI evaluation?

Yes, the findings suggest a shift toward live, consequence-based testing that emphasizes management skills, which could become part of future AI certification and assessment frameworks.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Build a Recovery Plan for Stolen Devices

Securing your stolen devices with a solid recovery plan is crucial—discover essential steps to protect your data and increase your chances of recovery.

AI output review queue for customer support macros

Support teams are testing a new AI macro review queue to ensure policy compliance and tone accuracy before deployment.

Vendor insurance certificate tracker for property managers

Small property managers are set to trial a new vendor insurance certificate tracker to streamline document management and risk control.

The High-End PC and Workstation Tax

Memory costs surge in 2026, making DIY PC building more expensive and shifting the market dynamics for high-end builds and workstations.