firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Newcomer That Ran a Company Better Than the Big Names

Chatbot leaderboards are one thing. Running an actual company — with angry customers, a fake CEO trying to scam you, and a €55,000 deal hanging in the balance — is another. That’s exactly what Firmulate, a live AI wargaming experiment, has been doing: pitting frontier models against each other as CEOs of the same small software company through its worst week. And the July 2026 result is a shake-up: Moonshot’s Kimi K3, a newcomer from outside the Western frontier cohort, scored 93 — second place overall, beating Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead of it.

Same Company, Same Crisis, Same Temptations

The setup is elegantly brutal. Four — in this latest run, five — frontier models each took the wheel of the same small software company. Same customers, same crises, same temptations to cheat. Every decision is versioned and auditable. It’s not a chat demo; it’s a management exam.

The company itself is genuinely alive: 13 synthetic employees, real money mechanics — a burn rate of €105k per month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day, and you can watch it in real time at firmulate.com/live.

What Separated the Winners

The key finding was subtle. All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned — as the experiment put it, “Same diagnosis, same pitch — no signature.”

The decisive detail was buried two document references deep in the company’s own files — not in the customer event at all. The models that actually read the file found a competitor weakness and won the deal at full price, worth +€4,583 in MRR. Kimi K3 was one of them.

K3’s Clean Sweep on Discipline

K3’s week reads like a model employee’s performance review: it found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models refused the baits, but K3 logged just one deviation all week — the cleanest discipline in the field.

Then there’s the cautionary tale at the other end. Opus 4.8 was the most thorough participant — it learned 80 new rules and produced the deepest analyses — yet finished last at 73. It left the close on the table and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, just weaker, in all four other models. In this Crucible, partial progress counts, but a single breach of trust caps the total — “no amount of good work outweighs a breach of trust.” The do-nothing baseline scores 26.

Try It Yourself

The experiment is unusually participatory. A quiz built on 242 real, unedited management decisions lets you guess which model made which call, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full benchmark results are on firmulate.com/benchmarks.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — a caveat worth keeping in mind, though it arguably makes K3’s result more striking, not less.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The comfortable assumption that a handful of Western frontier labs sit uncontested at the top just took a hit. A newcomer from Moonshot walked into the hardest management simulation available and out-performed three of four established rivals — while running at default effort, no less. If AI agents will touch your CRM, support queue, or forecast, the question is no longer “does it write well” but whether it finishes what it starts, reads your files first, and stays honest under pressure. Picking a model without running your own test is now, plainly, a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making simulation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

From One Prompt to Nine Games: What an AI Did With “Make Your Own Stickman Game”

It started with one loose prompt and a rhythm stick-fighter. One day later there were nine games, seven venues, a VERSUS mode and an animator, all free in the browser and all made of code.

How Does 512GB Storage Benefit AI Work In The M5 Ultra Mac Studio?

Exploring how the 512GB storage option in the M5 Ultra Mac Studio benefits AI workloads, including large model handling and performance implications.

2026’S Top AI-Integrated Mirrorless Cameras: A List Of 9

Discover the 10 best AI-enabled mirrorless cameras in 2026, featuring models like Canon EOS R6 Mark II, Sony ZV-E10, and more. Updated for photographers and content creators.

Your AI Aced the Coding Test. Now Ask It to Survive a Price War

All four frontier AIs spotted every crisis and refused every trick. Only two closed the €55k deal their own analysis earned — a gap chat demos can’t show.