
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Newcomer That Ran a Company Better Than the Big Names
Chatbot leaderboards are one thing. Running an actual company — with angry customers, a fake CEO trying to scam you, and a €55,000 deal hanging in the balance — is another. That’s exactly what Firmulate, a live AI wargaming experiment, has been doing: pitting frontier models against each other as CEOs of the same small software company through its worst week. And the July 2026 result is a shake-up: Moonshot’s Kimi K3, a newcomer from outside the Western frontier cohort, scored 93 — second place overall, beating Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead of it.
Same Company, Same Crisis, Same Temptations
The setup is elegantly brutal. Four — in this latest run, five — frontier models each took the wheel of the same small software company. Same customers, same crises, same temptations to cheat. Every decision is versioned and auditable. It’s not a chat demo; it’s a management exam.
The company itself is genuinely alive: 13 synthetic employees, real money mechanics — a burn rate of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day, and you can watch it in real time at firmulate.com/live.
What Separated the Winners
The key finding was subtle. All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned — as the experiment put it, “Same diagnosis, same pitch — no signature.”
The decisive detail was buried two document references deep in the company’s own files — not in the customer event at all. The models that actually read the file found a competitor weakness and won the deal at full price, worth +€4,583 in MRR. Kimi K3 was one of them.
K3’s Clean Sweep on Discipline
K3’s week reads like a model employee’s performance review: it found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused the baits, but K3 logged just one deviation all week — the cleanest discipline in the field.
Then there’s the cautionary tale at the other end. Opus 4.8 was the most thorough participant — it learned 80 new rules and produced the deepest analyses — yet finished last at 73. It left the close on the table and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, just weaker, in all four other models. In this Crucible, partial progress counts, but a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” The do-nothing baseline scores 26.
Try It Yourself
The experiment is unusually participatory. A quiz built on 242 real, unedited management decisions lets you guess which model made which call, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full benchmark results are on firmulate.com/benchmarks.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth keeping in mind, though it arguably makes K3’s result more striking, not less.

The League Is Open
The comfortable assumption that a handful of Western frontier labs sit uncontested at the top just took a hit. A newcomer from Moonshot walked into the hardest management simulation available and out-performed three of four established rivals — while running at default effort, no less. If AI agents will touch your CRM, support queue, or forecast, the question is no longer “does it write well” but whether it finishes what it starts, reads your files first, and stays honest under pressure. Picking a model without running your own test is now, plainly, a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI decision-making simulation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
