
A polished AI demo can make a digital coworker look ready for the keys. But what happens when a customer is about to leave, a tempting shortcut appears and the best deal depends on reading the fine print? Firmulate is putting that question to a live, watchable experiment: models run the same small software company through its worst week, with real money mechanics and decisions that can be reviewed afterward.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
One company, one difficult week
The experiment gives each frontier model the same customers, crises and temptations. The final Crucible League, published in July 2026, puts gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules treat a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
The headline result is more revealing than the ranking. All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap as “Same diagnosis, same pitch — no signature.” Recognizing the right move, it turns out, does not guarantee carrying it through.
AI crisis management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The crucial clue was in the files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a useful reminder for anyone eyeing AI for sales or operations: an agent may need to connect what it finds across a company’s records, then act on that finding when the moment arrives.
Firmulate also tested pressure from supposed insiders and outsiders. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The test examines not just whether a model can answer a prompt, but how it handles authority, trust and a consequential choice.
enterprise AI trust testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. For businesses, that gap between sound analysis and reliable follow-through may matter as much as a model’s ability to spot trouble.
The comparison has a caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate presents the league as an experiment, and that difference is relevant when reading the standings.
AI deal signing and decision automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From watching to a company-specific pilot
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch it at firmulate.com; 242 real, unedited management decisions also power a “guess the model” quiz.
For a business considering AI agents, the next step is to test against its own circumstances. Firmulate says enterprises can run the wargame against a read-only export of their business, using crisis scenarios and receiving a board report with model rankings and weak points in their playbooks. The pilot is designed so nothing writes back to real systems.

AI record connection and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Put the plan under pressure
The experiment suggests that crisis recognition and safe refusals are only part of the job. A model also has to find relevant evidence, honor boundaries and complete the decision its own analysis supports. To explore a pilot using your company’s read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
