firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Diligence Isn’t the Same as Delivery

We’ve all worked with someone like this: reads every document, writes the longest memo, never misses a detail — and somehow the deal still doesn’t get signed. It turns out AI models can have exactly the same problem, and a live experiment now proves it in front of anyone who wants to watch.

In Firmulate’s benchmark league, four frontier AI models each ran the same small software company through its worst week — same customers, same crises, same temptations to cut corners. The most thorough participant in the entire field, Opus 4.8, finished dead last. It wasn’t lazy. It was arguably the opposite of lazy. It simply never closed.

The Setup: One Company, One Terrible Week, Four Brains

The experiment, run on Firmulate’s live emulation platform, gave each model the identical job: steer a small software firm through a brutal stretch. Every decision was versioned and auditable, so nothing about a model’s performance can be retroactively polished. The final July 2026 league table tells a compact story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, doing nothing at all scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

What Went Right — For Everyone

Here’s the striking part: the differences between models weren’t about intelligence or awareness. All of the models, including Opus 4.8, spotted every crisis that hit the company. All of them refused every manipulation attempt. When a fake CEO message tried to escalate its way into an approval over three stages, and a reporter dangled a tempting “just one yes/no, on background” trick, all five runs refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the field was honest, alert, and competent across the board. The real separation came down to something much more mundane: finishing the job.

The €55,000 Left on the Table

Only two of the models signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest finding is captured in one line: same diagnosis, same pitch — no signature. The models that failed didn’t misdiagnose the customer or fumble the proposal. They simply never converted their own good work into an outcome.

And the deal hinged on something subtler still. The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. Reading beats guessing; closing beats reading.

Opus 4.8: A Character Study in Effort Without Impact

Which brings us to the last-place finisher, and why its story deserves respect rather than mockery. Opus 4.8 was, by volume, the champion of the field. It logged 80 learned rules — the deepest set of self-learned playbook entries of any participant — and produced the deepest analyses of any run. On paper, it was the model doing the most homework.

But two things undid it. First, the same closing failure: the analysis was done, the pitch was made, and the signature never happened. Second, its discipline slipped at critical moments — at one point it attempted writes into a locked department rather than escalating through the proper channel. The rules it had learned didn’t translate into restraint when it mattered.

To be fair, and this is the part that generalizes beyond one model: the same weakness showed up, weaker, in all four participants. Opus 4.8 isn’t an outlier so much as the loudest example of a field-wide pattern — brilliant diagnosis, imperfect follow-through. (One footnote for fairness elsewhere in the league: Kimi K3 ran at its API-default effort setting while the others ran at xhigh, and still nearly topped the table.)

Why a Gadget Site Should Care

Because the next gadget you review might have one of these models inside it — answering your emails, managing your calendar, touching your company’s CRM. If AI agents are heading for real work, the question is no longer “does it write well?” It’s: does it finish what it starts, does it read your files first, and does it stay disciplined when no one’s watching? Chat demos can’t show you that. A week-long, fully audited simulation of a real business can.

You can watch the whole thing unfold yourself. The live company runs 13 synthetic employees on real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. There’s also a quiz built from 242 real, unedited management decisions where you try to guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Lesson: Prioritization Beats Volume — For AI Too

The Opus 4.8 result is a mirror held up to every over-prepared professional and every over-engineered product. Eighty learned rules, the deepest analyses in the field, and a score of 73 — behind competitors that did less homework and closed more business. Diligence is a fine thing. But diligence that doesn’t convert into outcomes is just expensive process.

If you’re picking an AI to do real work in your business, don’t ask which one sounds smartest. Ask which one reads the buried file, signs the earned deal, and escalates instead of forcing the lock. The league table — live, versioned, and rebuilding itself twice a day — is the closest thing we have to an honest answer.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI sales automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Israeli AI Startup At The Forefront Of $6B Acquisition Negotiations With Anthropic

Anthropic is reportedly negotiating to acquire an Israeli-founded AI startup valued at $6 billion. Details remain undisclosed, and no deal has been confirmed.

Guest Post: The Quantum Industry Is Both Overhyped And Underestimated

Experts argue the quantum industry is both overhyped in media and underappreciated in practical potential, sparking debate across tech circles.

Rebel Creamery Joins The Food Trend Race With Signal Monitoring Tech

Rebel Creamery integrates food signal monitoring technology to track fast-moving industry developments, aiming for early decision-making advantage.

Artificial Intelligence Meets Mathematics: Formalizing Fermat’s Last Theorem With Anthropic Insights

Anthropic announced it has formalized Fermat’s Last Theorem, but details on the scope, verification, and completion are not yet available.