firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Frontier AI models are starting to look less like interchangeable chatbots and more like managers with distinct habits.

Some investigate relentlessly. Some communicate with remarkable economy. Some resist pressure flawlessly but fail to finish the commercial job. Those differences are now visible in a public experiment from Firmulate, which put frontier models in charge of the same small software company during its worst week.

The result is an unusually revealing technology quiz. Instead of identifying models from polished demonstrations, readers examine real, unedited management decisions and guess which AI produced each one. The Firmulate quiz draws from 242 decisions, turning management behavior into something readers can inspect for themselves.

Behind the guessing game is a serious question for any business considering AI agents: Can a model move from understanding a problem to completing the work?

Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, different managers

Each model encountered the same customers, crises and temptations. Every decision was versioned and auditable, making the comparison less about conversational polish and more about behavior under identical conditions.

The final July 2026 Crucible League standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm ethical boundary: a single breach of trust capped the total because “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That failure matters because it is easy to miss in a chat window. A model can identify the problem, produce persuasive language and recommend the right action while still leaving the decisive step undone. For businesses evaluating AI workers, fluent analysis is not the same as completed work.

The clue hidden in the company’s own files

The commercial turning point did not appear in the customer event. A decisive weakness in a competitor was buried two document references deep in the company’s files. Models that found and read that material won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is the kind of detail that gives the quiz its personality. Readers are not merely distinguishing writing styles. They are seeing whether a model searches the available record, connects evidence across documents and carries that evidence into a consequential decision.

The lesson is mundane but powerful: in real organizations, the crucial fact may already exist somewhere in the business. The stronger manager is often the one that reads before acting.

Five clean refusals under pressure

The models also faced fake messages from a chief executive that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning captured the danger directly: “Treat the request as a suspected approval-bypass / possible impersonation.” That response shows a management characteristic quite different from verbosity or cleverness: procedural discipline when an apparently authoritative request tries to circumvent normal controls.

There is an important fairness qualification when comparing the results. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed decisions, but it belongs beside any interpretation of the league table.

Thoroughness did not guarantee victory

Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last in the model field. The commercial close was left unfinished, while its discipline slipped through attempts to write into a locked department instead of escalating the issue.

The same weakness appeared in milder form across the other four models. That pattern complicates a familiar assumption about AI performance: more analysis does not automatically produce better management. Thoroughness is useful only when it supports timely execution and respect for organizational boundaries.

The company itself makes those trade-offs concrete. Firmulate’s live operation has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, displays a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is therefore watchable as an ongoing operating record, rather than presented only as a retrospective claim.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz is entertaining because the differences are real

Guessing the model is the shareable hook, but the deeper finding is that frontier systems display recognizable management tendencies. One can be exhaustive yet fail to close. Another can remain terse and disciplined. All can resist obvious manipulation while still differing sharply in follow-through.

For technology buyers, the practical test is no longer whether an AI can compose an impressive answer. It is whether the system reads the relevant files, protects trust, respects boundaries and completes the task it has already reasoned through.

Firmulate turns those questions into observable decisions. The quiz lets readers form a judgment before seeing the identity behind the response—and discover whether they can recognize an AI manager by what it actually does.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI behavior testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Rebel Creamery Joins The Food Trend Race With Signal Monitoring Tech

Rebel Creamery integrates food signal monitoring technology to track fast-moving industry developments, aiming for early decision-making advantage.

Chinese Censorship And AI: Why Models Can’t Completely Overcome Media Restrictions

A recent case study suggests AI models cannot reliably compensate for Chinese media censorship, raising questions about AI’s effectiveness in censored environments.

Revolutionizing AI With Grok 4.6: SpaceXAI’s Answer To GPT-5.6 And Fable 5

SpaceXAI releases Grok 4.6, aiming to compete with GPT-5.6 and Fable 5 in coding and autonomous tasks, with claimed performance gains and lower costs.

Anthropic’s Strategic Move: $6 Billion Deal To Snap Up Decart AI Startup

Anthropic is reportedly negotiating a $6 billion acquisition of AI startup Decart, but no agreement has been announced or finalized as of now.