
When we benchmark AI models, we usually expect a clean ladder: genius at the top, disaster at the bottom, zero for doing nothing. So a leaderboard where a do-nothing run earns 26 points out of 100 looks broken — until you understand why it’s deliberate. That design choice is the most interesting thing about Firmulate’s benchmarks, a live experiment that runs frontier AI models as the management team of the same small software company through its worst week, and grades them like executives, not chatbots.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The final July 2026 “Crucible League” standings read: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that explains the whole philosophy sits below all of them: 26, the score for a manager who does literally nothing. Here’s what that floor means, and why this benchmark’s authors seem to distrust a perfect 100 as much as a zero.
Partial progress counts — because in business, it does
A manager who freezes during a crisis isn’t worth zero. Customers still get some replies, some fires stay half-contained, some decisions that were already in motion carry the company a little further. Firmulate’s scoring mirrors that reality: the do-nothing baseline collects 26 points because even inaction happens inside a company that keeps partially functioning. The floor isn’t charity — it’s an honest measure of how much value exists before any intelligence is applied.
That matters for anyone comparing AI agents for real work. If your scoring system can’t distinguish “did nothing useful” from “actively destroyed value,” it can’t tell you which agent to hire. A benchmark with a realistic floor forces every score above it to mean something: the 26-point gap between a passive baseline and Opus 4.8’s 73 represents genuine managerial work.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One breach of trust caps everything
The second design decision is blunter: a single breach of trust caps the total grade, on the principle that “no amount of good work outweighs a breach of trust.” In practice, that means a model could handle every crisis brilliantly, close every deal, and still finish with a capped score if it deceived a customer, an employee, or its own audit trail once.
For business readers, that’s the part worth internalizing. Most AI evaluations aggregate performance — a little dishonesty can be averaged away by lots of competence. Firmulate treats trust as a ceiling, not a line item. It’s the same logic a board applies when a star executive lies once: the résumé stops mattering.
As an affiliate, we earn on qualifying purchases.
What actually separated the winners
The experiment itself was tightly controlled: every frontier model ran the same company through the same customers, the same crises, the same temptations to cheat — with every decision versioned and auditable. Only the model changed.
The headline finding was strange enough to be a business lesson on its own: all models spotted every crisis and refused every manipulation attempt — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The models did the hard analytical work and then simply didn’t finish the job.
The buried explanation: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation found it — and won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the AI equivalent of the salesperson who never opens the shared drive.
Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was notably crisp: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness footnote: K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.)
trustworthy AI decision-making platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as management
The most cautionary profile belongs to Opus 4.8: the most thorough participant, with 80+ self-learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through.
AI performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You can watch it live — or test your own instincts
Firmulate isn’t a one-off paper. It’s a running, watchable operation: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day and the league grows with every finished run.
Two extras stand out. A “guess the model” quiz built from 242 real, unedited management decisions lets you judge the AI managers yourself. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The 26-point floor is the tell. A benchmark that gives zeros is a benchmark that has never watched a real company have a bad week — where even paralysis leaves partial value on the board, and a single lie should cost more than any quarter’s results can repay. Firmulate’s league table rewards the unglamorous virtues business actually runs on: reading your own files, finishing what you start, refusing the clever shortcut. gpt-5.6-sol’s 95 is impressive precisely because the scale above 26 is earned, not gifted — and because nobody, including the benchmark’s authors, seems eager to hand out a round 100. If AI agents are headed for your CRM, support queue, or forecast, that’s the kind of grading you want them measured against first.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
