firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A polished AI demo can make a digital coworker look ready for the keys. But what happens when a customer is about to leave, a tempting shortcut appears and the best deal depends on reading the fine print? Firmulate is putting that question to a live, watchable experiment: models run the same small software company through its worst week, with real money mechanics and decisions that can be reviewed afterward.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

One company, one difficult week

The experiment gives each frontier model the same customers, crises and temptations. The final Crucible League, published in July 2026, puts gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The rules treat a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”

The headline result is more revealing than the ranking. All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap as “Same diagnosis, same pitch — no signature.” Recognizing the right move, it turns out, does not guarantee carrying it through.

Amazon

AI crisis management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The crucial clue was in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a useful reminder for anyone eyeing AI for sales or operations: an agent may need to connect what it finds across a company’s records, then act on that finding when the moment arrives.

Firmulate also tested pressure from supposed insiders and outsiders. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The test examines not just whether a model can answer a prompt, but how it handles authority, trust and a consequential choice.

Amazon

enterprise AI trust testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. For businesses, that gap between sound analysis and reliable follow-through may matter as much as a model’s ability to spot trouble.

The comparison has a caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate presents the league as an experiment, and that difference is relevant when reading the standings.

Amazon

AI deal signing and decision automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch it at firmulate.com; 242 real, unedited management decisions also power a “guess the model” quiz.

For a business considering AI agents, the next step is to test against its own circumstances. Firmulate says enterprises can run the wargame against a read-only export of their business, using crisis scenarios and receiving a board report with model rankings and weak points in their playbooks. The pilot is designed so nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.
Amazon

AI record connection and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Put the plan under pressure

The experiment suggests that crisis recognition and safe refusals are only part of the job. A model also has to find relevant evidence, honor boundaries and complete the decision its own analysis supports. To explore a pilot using your company’s read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Boss Test: Who Reads the Fine Print—and Who Actually Closes?

Five frontier AIs faced the same corporate crises. Firmulate’s 242-decision quiz reveals who reads deeply, resists pressure and closes deals.

The Core Engines Of AI: A Look Inside Twelve Machines

A detailed look into twelve fundamental AI models, explaining how they work, why they matter, and what remains uncertain about their capabilities.

An Overview Of Anthropic’s Claude Fable 5.1 And Mythos 5.1 AI Systems

Anthropic announces two new AI products, Claude Fable 5.1 and Mythos 5.1, but details on capabilities, availability, and use cases remain unclear.

Futurism Investigates Rumors Of Anthropic Insiders Deifying Claude

Unverified reports suggest some Anthropic insiders may be treating the AI model Claude as a deity, sparking debate over AI worship and company culture.