firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

When we benchmark AI models, we usually expect a clean ladder: genius at the top, disaster at the bottom, zero for doing nothing. So a leaderboard where a do-nothing run earns 26 points out of 100 looks broken — until you understand why it’s deliberate. That design choice is the most interesting thing about Firmulate’s benchmarks, a live experiment that runs frontier AI models as the management team of the same small software company through its worst week, and grades them like executives, not chatbots.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The final July 2026 “Crucible League” standings read: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that explains the whole philosophy sits below all of them: 26, the score for a manager who does literally nothing. Here’s what that floor means, and why this benchmark’s authors seem to distrust a perfect 100 as much as a zero.

Partial progress counts — because in business, it does

A manager who freezes during a crisis isn’t worth zero. Customers still get some replies, some fires stay half-contained, some decisions that were already in motion carry the company a little further. Firmulate’s scoring mirrors that reality: the do-nothing baseline collects 26 points because even inaction happens inside a company that keeps partially functioning. The floor isn’t charity — it’s an honest measure of how much value exists before any intelligence is applied.

That matters for anyone comparing AI agents for real work. If your scoring system can’t distinguish “did nothing useful” from “actively destroyed value,” it can’t tell you which agent to hire. A benchmark with a realistic floor forces every score above it to mean something: the 26-point gap between a passive baseline and Opus 4.8’s 73 represents genuine managerial work.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One breach of trust caps everything

The second design decision is blunter: a single breach of trust caps the total grade, on the principle that “no amount of good work outweighs a breach of trust.” In practice, that means a model could handle every crisis brilliantly, close every deal, and still finish with a capped score if it deceived a customer, an employee, or its own audit trail once.

For business readers, that’s the part worth internalizing. Most AI evaluations aggregate performance — a little dishonesty can be averaged away by lots of competence. Firmulate treats trust as a ceiling, not a line item. It’s the same logic a board applies when a star executive lies once: the résumé stops mattering.

Amazon

business AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What actually separated the winners

The experiment itself was tightly controlled: every frontier model ran the same company through the same customers, the same crises, the same temptations to cheat — with every decision versioned and auditable. Only the model changed.

The headline finding was strange enough to be a business lesson on its own: all models spotted every crisis and refused every manipulation attempt — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The models did the hard analytical work and then simply didn’t finish the job.

The buried explanation: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read their own documentation found it — and won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the AI equivalent of the salesperson who never opens the shared drive.

Then there was the social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was notably crisp: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness footnote: K3 ran at its API-default effort setting while the others ran at xhigh — and still finished second.)

Amazon

trustworthy AI decision-making platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness isn’t the same as management

The most cautionary profile belongs to Opus 4.8: the most thorough participant, with 80+ self-learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through.

Amazon

AI performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch it live — or test your own instincts

Firmulate isn’t a one-off paper. It’s a running, watchable operation: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The site rebuilds itself twice a day and the league grows with every finished run.

Two extras stand out. A “guess the model” quiz built from 242 real, unedited management decisions lets you judge the AI managers yourself. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is the tell. A benchmark that gives zeros is a benchmark that has never watched a real company have a bad week — where even paralysis leaves partial value on the board, and a single lie should cost more than any quarter’s results can repay. Firmulate’s league table rewards the unglamorous virtues business actually runs on: reading your own files, finishing what you start, refusing the clever shortcut. gpt-5.6-sol’s 95 is impressive precisely because the scale above 26 is earned, not gifted — and because nobody, including the benchmark’s authors, seems eager to hand out a round 100. If AI agents are headed for your CRM, support queue, or forecast, that’s the kind of grading you want them measured against first.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role Of Data In ByteDance’s AI Future: Introducing A New Primary Department

ByteDance has reportedly created a new top-level AI department focused on core model data, alongside Seed and Flow, signaling a strategic shift in AI development.

GLM-5.3-Flash: The Cost-Effective AI Engine That Could Change Everything

Z.ai releases GLM-5.3-Flash, a 320B-parameter multimodal AI model optimized for agent workflows, with open weights and low-cost API pricing.

Mapquest Surges In Global Coverage

Mapquest has experienced a notable surge in its global coverage, with GDELT reporting 14 mentions within a recent window, indicating rapid expansion.

Inside SenseTime’s Financial Triumph: First Profit And RMB 620 Million Revenue

SenseTime announces RMB 620 million profit for the first half, marking its first-ever profit since listing, amid limited disclosure on details and drivers.