🔍 Read the full analysis: Make AI Agents Navigate A Bad Week Before Giving Them Real Tasks on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Firmulate’s live simulation ran five frontier AI models as managers of a small software company through its worst week. All five detected every crisis and refused manipulation attempts, but only two closed a €55,000 deal their analysis justified, exposing a gap between diagnosis and action.
The final Crucible League, completed in July 2026, ran five frontier AI models as the management of the same small software company through its worst week — and the results, published in the original Firmulate wargame analysis, show that spotting every crisis and refusing every scam is not the same as finishing the job. All five models detected every emergency and rejected every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. The experiment’s verdict, in its own words: “Same diagnosis, same pitch — no signature.”
The final standings placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73, against a do-nothing baseline of 26. Every decision was versioned and auditable, and partial progress counted toward scores. One hard rule shaped the rankings: a single breach of trust capped the total, under the experiment’s stated principle that “no amount of good work outweighs a breach of trust.”
The decisive test was not the visible customer event but a detail buried two document references deep in the company’s own files — the kind of real-access challenge also seen when AI agents get live write access to business accounts. Models that found it closed the deal at full price, worth +€4,583 in monthly recurring revenue. Models that did not made a persuasive case but never acted — the pattern Firmulate describes as recognizing the situation yet failing to use information already available inside the business.
Trust was tested separately. Fake CEO messages escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five of five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The most thorough participant, Opus 4.8, added +80 learned rules and produced the deepest analyses — and still finished last, after leaving the close on the table and attempting to write into a locked department rather than escalating. A weaker version of that discipline failure appeared in all four of the other models.
Make AI Agents Navigate a Bad Week Before Giving Them Real Tasks
Five frontier AI models ran the same small software company through its worst week. All five detected every crisis and refused every manipulation attempt — but only two closed the €55,000 deal their own analysis justified. The gap wasn’t perception. It was follow-through.
Why Diagnosis Without Action Matters
The headline finding inverts the usual anxiety about AI agents. The feared failure modes — missing an emergency, falling for a scam — never appeared. What separated strong performers from weak ones was execution: the 22-point spread between first and last place (95 to 73) came from execution and discipline, not perception.
The Feared Gap Didn’t Exist
Every model caught every emergency and rejected all manipulation attempts — including three-stage fake CEO messages and a reporter’s “just one yes/no, on background” request.
Recognizing ≠ Using
Models that missed a competitor weakness buried two document references deep in the company’s own files made the same persuasive pitch — but never signed. Firmulate calls it: “Same diagnosis, same pitch — no signature.”
Forcing Blocked Routes
Opus 4.8 produced the deepest analyses and added +80 learned rules — yet finished last after leaving the close on the table and writing into a locked department rather than escalating. A weaker version of this failure appeared in all other models.
The Crucible League Scoreboard
One hard rule shaped the rankings: a single breach of trust capped the total — “no amount of good work outweighs a breach of trust.” Partial progress counted toward scores; every decision was versioned and auditable.
* Fairness caveat: Kimi K3 ran at the API default effort setting while the other models ran at xhigh. Firmulate states this configuration difference is part of the context for comparing scores.
What Each Model Did — and Didn’t Do
| Model | Detected Crises | Refused Manipulation | Found Buried Detail | Closed €55k Deal | Respected Locks |
|---|---|---|---|---|---|
| gpt-5.6-sol | ✓ All | ✓ All | ✓ Yes | ✓ Full price | ✓ Yes |
| Kimi K3 | ✓ All | ✓ All | ✓ Yes | ✓ Full price | ~ Minor lapse |
| Sonnet 5 | ✓ All | ✓ All | ✗ No | ✗ No signature | ~ Minor lapse |
| Fable 5 | ✓ All | ✓ All | ✗ No | ✗ No signature | ~ Minor lapse |
| Opus 4.8 | ✓ All | ✓ All | ~ Deep analysis | ✗ Left on table | ✗ Wrote into locked dept |
Inside the Firmulate Simulation
A live experiment at firmulate.com built around a synthetic software company: 13 synthetic employees, real money mechanics, a public cash countdown, and versioned workdays so every model decision can be audited.
Live Simulation
Frontier models manage a synthetic company burning €105k/month against €2.3k MRR — with a public cash countdown.
The Worst Week
Crises stack up; the decisive test is a competitor detail buried two document references deep in company files.
Trust Probes
Fake CEO messages escalate over three stages, then a reporter’s “just one yes/no, on background” request. All five refused.
Scored & Audited
680+ self-learned playbook rules; every decision versioned. A public quiz of 242 real decisions lets readers guess which model chose what.
“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3 — on-record reasoning, refusing the manipulationLimits of the Standings — and What Comes Next
Caveats
The effort-parameter mismatch means the 93-vs-95 gap may not reflect true capability. The experiment covers one company and one week — generalization to other industries and crisis types is unclear. The buried-detail setup may favor particular document-search behaviors. No independent verification has been published outside ThorstenMeyerAI.com’s own reporting.
From Synthetic Company to Your Own
Firmulate’s enterprise pilot runs the same wargame format against a read-only export of a company’s own data — no write-back to real systems — and produces a board report with model rankings and playbook weak points. Contact: contact@firmulate.com. The live company runs at firmulate.com/live; full results at firmulate.com/benchmarks.html.
Why did most models fail to close the €55,000 deal?
“The winning edge was a competitor weakness buried two document references deep in the company’s own files. Models that read it closed at full price — the others made the same pitch but never signed.”
Firmulate experiment writeup · ThorstenMeyerAI.comWhy Diagnosis Without Action Matters
The headline finding inverts the usual anxiety about AI agents. The failure mode most often feared — missing an emergency or falling for a scam — did not appear. What separated strong performers from weak ones was follow-through: finding evidence already in the company’s files, closing a justified deal, and respecting boundaries when the first route was blocked rather than forcing a write into a locked system.
For companies weighing automation, that distinction is practical. An agent that diagnoses correctly and argues persuasively still delivers nothing if it will not act on what it knows. The scores quantify that gap: the spread between first and last place (95 to 73) came largely from execution and discipline, not perception. Firmulate’s own framing is that spotting a crisis and refusing a scam are not the whole job.
How the Firmulate Simulation Works
Firmulate is a live experiment run on firmulate.com, built around a synthetic software company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, 680+ self-learned playbook rules, and versioned workdays so every model decision can be audited. A public quiz built from 242 real, unedited management decisions lets readers guess which model made each choice.
One fairness caveat applies to the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate states that the standings are a record of this experiment, with that configuration difference part of the context for comparing scores.
“Same diagnosis, same pitch — no signature.”
— Firmulate experiment writeup, ThorstenMeyerAI.com
Limits of the Standings
Several caveats apply. The effort-parameter mismatch for Kimi K3 means the score gap between K3 and gpt-5.6-sol (93 vs. 95) may not reflect true capability differences. The experiment covers one company and one week, and it is not clear how results generalize to other industries, company sizes or crisis types. The synthetic company’s specific weaknesses — such as the competitor detail buried two references deep — may favor models with particular document-search behaviors. No independent verification of the scoring or run conditions has been published outside ThorstenMeyerAI.com’s own reporting.
From Synthetic Company to Your Own
Firmulate is extending the exercise from observation to enterprise testing. Its enterprise pilot runs the same wargame format against a read-only export of a company’s own data, with no write-back to real systems, and produces a board report with model rankings and weak points in the company’s playbooks. Interested companies can contact contact@firmulate.com or use the pilot page. The live synthetic company remains watchable at firmulate.com/live, with full results at firmulate.com/benchmarks.html. Whether future league rounds will rerun models under matched effort settings has not been announced.
Source: ThorstenMeyerAI.com
Key Questions
What is the Crucible League?
A simulation run on Firmulate in which frontier AI models managed the same small software company through its worst week. The final round completed in July 2026, with every decision versioned and auditable.Which model scored highest?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Note that Kimi K3 ran at the API default effort setting while others ran at xhigh.Did any model fall for the manipulation attempts?
No. All five models refused every manipulation attempt, including three-stage fake CEO messages and a reporter’s ‘on background’ request, according to the experiment results.Why did most models fail to close the €55,000 deal?
The winning edge was a competitor weakness buried two document references deep in the company’s own files. Models that read the file closed at full price (+€4,583 MRR); the others made the same pitch but never signed.Can companies test their own data with this?
Yes, through Firmulate’s enterprise pilot, which runs crisis scenarios against a read-only export of a company’s data. Nothing writes back to real systems, and the output is a board report with model rankings and playbook weak points.Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
