
Two AIs Closed a €55,000 Deal. The Other Three Choked on Paperwork.
When people compare AI models, they usually argue about witty answers and benchmark trivia. But a new kind of test asks a different question: can an AI agent actually finish a job — one that requires digging through a company’s own documents, spotting a decisive fact buried two references deep, and closing a deal at full price?
That’s the premise behind Firmulate’s benchmark league, which ran four frontier AI models through the same nightmare week running a small software company. All of them aced the obvious parts. Only some of them did their homework — and the difference was worth €55,000.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crisis, Different Brains
The setup is elegantly cruel: each model gets the identical small software firm, the same customers, the same crises, and the same temptations to cheat. Every decision is versioned and auditable, so nothing hides in the margins. The final July 2026 league table tells the story:
- gpt-5.6-sol — 95 points, the complete performance
- Kimi K3 — 93 points, the newcomer with the cleanest discipline in the field
- Fable 5 — 77 points
- Opus 4.8 — 73 points, last place despite being the most thorough participant
Sonnet 5 — 88 points
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
The Finding That Should Worry Every Vendor
Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment sums it up bluntly: “Same diagnosis, same pitch — no signature.”
Why did the others stall? The decisive competitor weakness — the fact that justified closing at full price — wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, lost it automatically.
In other words, “reads your files before answering” isn’t a nice-to-have chat feature. It’s a measurable, purchase-deciding property of AI agents — and it’s invisible in a standard demo.
The Social Engineering Gauntlet
The week also included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s exactly the instinct you want in anything touching your CRM or support queue.
Effort Isn’t Everything
The most surprising profile belongs to Opus 4.8: it learned the most rules (+80), produced the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. One fairness note: K3 ran at the API’s default effort setting while the others ran at xhigh — and still nearly won.
It’s Live, and You Can Play
This isn’t a one-off lab report. The company is running right now: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, watchable at firmulate.com/live. The site rebuilds itself twice a day as new runs finish.
There’s also a genuinely fun part: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

As an affiliate, we earn on qualifying purchases.
The Takeaway
Chat quality is solved. Finishing the job isn’t. The gap between a 95 and a 73 here wasn’t intelligence — it was whether the agent bothered to read the documents in front of it before acting. If you’re shopping for an AI workforce, don’t ask it to write you a poem. Ask it to find a fact two files deep and close the deal. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise AI document management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.