AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Why The AI Leaderboard After The Demo Is A Game Changer on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live experiment by Firmulate tested AI models managing a small company during its worst week. Results show management skills, not just chat quality, are key to effective AI performance. This could change how AI tools are assessed in business contexts.

In a groundbreaking live experiment, Firmulate tested five AI management models by having them run a small software company during its most challenging week, revealing that management quality, not just chat or technical responses, is crucial for success. The results suggest a shift in AI evaluation, emphasizing decision-making and trustworthiness in real-world business scenarios, as detailed in the original analysis.

The experiment involved five AI models competing in a simulated crisis environment, with the final July 2026 Crucible League ranking GPT-5.6-SOL first at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. For more on AI benchmarking, see the original analysis. A baseline scored only 26, underscoring the challenge of effective management under pressure.

Unlike traditional benchmarks that measure response quality or coding ability, this live test evaluated whether models could diagnose issues, communicate effectively, make decisions, and maintain trust. All models identified crises and rejected manipulation attempts, but only two secured the deal worth €55,000, highlighting that recognition alone isn’t enough—execution and trustworthiness matter more. Learn more about effective AI evaluation at the original analysis.

For instance, models that read relevant files and retrieved key facts closed deals at full price, while others failed to connect diagnosis with the final decision. The experiment also tested safety against social engineering, with all models refusing manipulative requests, but some still faltered in completing managerial tasks, exposing gaps between social compliance and effective action.

Interestingly, the most thorough model, Opus 4.8, performed poorly in final management execution despite deep analysis, illustrating that effort and activity do not necessarily translate into successful management. The results challenge the assumption that more effort or detailed analysis equals better performance, emphasizing the importance of disciplined decision-making and trust.

At a glance
reportWhen: ongoing, with final results announced i…
The developmentFirmulate’s live management test with AI models during a simulated crisis highlights the importance of managerial judgment over simple response quality, signaling a new approach to AI evaluation.

Transforming AI Evaluation Through Management Performance

This experiment demonstrates that AI’s ability to manage real-world business tasks—diagnosing issues, making decisions, maintaining trust—is more relevant than traditional benchmarks focused on technical or conversational prowess. It signals a potential shift in how organizations assess AI tools, prioritizing management quality and reliability for operational success.

For companies deploying AI, this means moving beyond response quality tests to live management scenarios that reveal whether models can handle complex, consequential tasks without compromising trust or effectiveness. Such an approach could redefine standards in AI evaluation and adoption.

Amazon

AI management tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Business

Current AI assessments often rely on coding challenges, chat arena responses, or partial progress metrics, which do not reflect real-world management complexities. The Firmulate experiment addresses this gap by testing models in a live, simulated company environment, where decisions have tangible consequences and trust is paramount.

Past benchmarks have focused on technical output or conversational fluency, but these do not capture the nuanced skills needed for effective management—such as prioritization, reading organizational context, and resisting shortcuts. The live test exposes these shortcomings, emphasizing the need for new evaluation paradigms.

Since July 2026, the experiment has highlighted that AI models can perform well on traditional metrics but still fail in critical management tasks, like closing deals or escalating issues appropriately, revealing an essential dimension often overlooked in AI development.

“The real test of AI management is whether it can handle consequences, prioritize effectively, and maintain trust under pressure.”

— Thorsten Meyer, lead researcher at Firmulate

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term Applicability

It remains unclear how well these findings generalize beyond the specific simulated environment used in the Firmulate experiment. The scalability of AI management performance in larger, more complex organizations is still to be tested. Additionally, the impact of different AI models, varying levels of training, and organizational contexts on management quality requires further exploration.

There is also uncertainty about how organizations will integrate such live management tests into their regular evaluation processes and whether this approach can be standardized across industries.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation Standards

Following these results, firms and AI developers are likely to explore more live, consequence-based testing environments, integrating management scenarios into AI evaluation frameworks. Future research may focus on refining metrics for trustworthiness, decision quality, and operational resilience.

Organizations considering AI for critical management tasks should begin assessing models in simulated or controlled environments that mimic real business pressures. The industry will also watch for broader adoption of these testing paradigms, potentially leading to new certification standards for AI management capabilities.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management quality more important than chat responses in AI evaluation?

Management quality reflects an AI’s ability to diagnose, decide, and maintain trust in real-world scenarios, which are crucial for operational success. Chat responses alone do not demonstrate these complex skills.

How does this experiment differ from traditional AI benchmarks?

Traditional benchmarks measure technical output or conversational ability, whereas this live experiment tests AI models’ capacity to manage a simulated company, focusing on decision-making, trust, and handling consequences.

What are the limitations of the current findings?

The experiment was conducted in a controlled, simulated environment, and its results may not fully translate to larger or more complex organizations. Further testing is needed to validate long-term applicability.

Could this lead to new standards for AI evaluation?

Yes, the findings suggest a shift toward live, consequence-based testing that emphasizes management skills, which could become part of future AI certification and assessment frameworks.

Source: ThorstenMeyerAI.com

You May Also Like

Japan’s Toto to invest $495m in chip materials, targeting 1-nm era

Japanese bathroom fixture maker Toto will invest $495 million over five years to expand into advanced semiconductor materials targeting 1-nanometer chip technology.

7 Best Gaming Laptop Prime Day Deals for 2026

Discover the best gaming laptop deals for Prime Day 2026, including the MSI Katana 17, Lenovo Legion Pro 7i, and more, with insights on discounts and value.

Create a Recovery USB Before You Need It

A well-prepared recovery USB can save your system—discover the essential steps to create one before emergencies strike.

How to Turn Off Battery-Draining Background Permissions

I can help you turn off battery-draining background permissions, but discover the simple steps to maximize your device’s battery life.