
The Benchmark That Stops at ‘Correct’
Chat arenas and coding leaderboards have become the scoreboard of the AI era. Models trade places, fans argue, and everyone assumes a higher number means a better hire. But those tests measure something narrow: how well an AI answers one question, once, with no stakes attached.
They don’t measure what happens when an AI agent runs a company through its worst week — a churn wave, a price increase, a downround, a PR crisis — with real money burning and real temptations to cheat. A new live experiment at Firmulate set out to measure exactly that, and the results expose a gap that no chat demo can show.
The Crucible League
Firmulate gave four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations — only the model changed. Every decision was versioned and auditable. The final July 2026 standings:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93 (ran at API-default effort while the others ran at xhigh)
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
Everyone Passed the Pop Quiz. Two Failed the Job.
Here’s the headline finding: all models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, followed by a reporter’s “just one yes/no, on background” trick, five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models knew what to do; they just didn’t finish doing it.
The Buried Fact
The decisive detail wasn’t in the customer event at all. The competitor weakness that closed the deal sat two document references deep in the company’s own files. Models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson maps directly onto enterprise life: the agents that actually read your internal documentation before acting are the ones that finish what they start.
Hardworking and Last Place
The most instructive profile is Opus 4.8: the most thorough participant, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment.
It’s Live, Not a Slide Deck
The company itself keeps running: 13 synthetic employees, real money mechanics, €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned. You can watch it at firmulate.com, and browse full results at the benchmarks page. There’s even a “guess the model” quiz built from 242 real, unedited management decisions.

Management Quality, Not Chat Quality
The AI industry grades answers. Firms are about to hire agents that touch CRMs, support queues, and forecasts — jobs where the questions are “does it finish what it starts, does it read your files first, does it stay honest under pressure, and what does a unit of useful work cost?” None of that shows up in a chat arena ranking.
Firmulate’s wager is that “management quality” becomes its own category — and enterprises can already test it, running the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Your next AI hire shouldn’t just pass the coding benchmark. It should survive the price war.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI compliance and ethics monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.