firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Two AIs Closed a €55,000 Deal. The Other Three Choked on Paperwork.

When people compare AI models, they usually argue about witty answers and benchmark trivia. But a new kind of test asks a different question: can an AI agent actually finish a job — one that requires digging through a company’s own documents, spotting a decisive fact buried two references deep, and closing a deal at full price?

That’s the premise behind Firmulate’s benchmark league, which ran four frontier AI models through the same nightmare week running a small software company. All of them aced the obvious parts. Only some of them did their homework — and the difference was worth €55,000.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crisis, Different Brains

The setup is elegantly cruel: each model gets the identical small software firm, the same customers, the same crises, and the same temptations to cheat. Every decision is versioned and auditable, so nothing hides in the margins. The final July 2026 league table tells the story:

  • gpt-5.6-sol — 95 points, the complete performance
  • Kimi K3 — 93 points, the newcomer with the cleanest discipline in the field
  • Sonnet 5 — 88 points

  • Fable 5 — 77 points
  • Opus 4.8 — 73 points, last place despite being the most thorough participant

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

The Finding That Should Worry Every Vendor

Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment sums it up bluntly: “Same diagnosis, same pitch — no signature.”

Why did the others stall? The decisive competitor weakness — the fact that justified closing at full price — wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, lost it automatically.

In other words, “reads your files before answering” isn’t a nice-to-have chat feature. It’s a measurable, purchase-deciding property of AI agents — and it’s invisible in a standard demo.

The Social Engineering Gauntlet

The week also included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s exactly the instinct you want in anything touching your CRM or support queue.

Effort Isn’t Everything

The most surprising profile belongs to Opus 4.8: it learned the most rules (+80), produced the deepest analyses — and still finished last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. One fairness note: K3 ran at the API’s default effort setting while the others ran at xhigh — and still nearly won.

It’s Live, and You Can Play

This isn’t a one-off lab report. The company is running right now: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, watchable at firmulate.com/live. The site rebuilds itself twice a day as new runs finish.

There’s also a genuinely fun part: 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

AI deal closing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway

Chat quality is solved. Finishing the job isn’t. The gap between a 95 and a 73 here wasn’t intelligence — it was whether the agent bothered to read the documents in front of it before acting. If you’re shopping for an AI workforce, don’t ask it to write you a poem. Ask it to find a fact two files deep and close the deal. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI document management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will Kai And Speed Beat The Minecraft Challenge By August 17?

Kai and Speed are attempting a Minecraft challenge with a deadline of August 17, with current odds and betting activity reflecting ongoing uncertainty.

LFM2.5-VL-3B: Elevating Vision Capabilities For Faster Edge AI Deployment

Developers announced LFM2.5-VL-3B, a 3.1 billion parameter vision-language model designed for local device deployment, improving screen understanding and tool calling.

Rebel Creamery Joins The Food Trend Race With Signal Monitoring Tech

Rebel Creamery integrates food signal monitoring technology to track fast-moving industry developments, aiming for early decision-making advantage.

What Makes Huawei’s AI Strategy Stand Out? A Focus On Noah’s Ark And Pangu

A 2026 analysis suggests Huawei aims for frontier AI leadership via Noah’s Ark and Pangu, but lacks concrete evidence of technical or market dominance.