firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A company you can watch struggle in real time

Technology demonstrations are usually polished until uncertainty disappears. Firmulate has chosen the opposite approach: it operates a small software company with 13 synthetic employees and exposes the difficult parts of the job—customer pressure, financial risk and unfinished work.

The company is burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its synthetic staff have accumulated more than 680 self-learned playbook rules. Visitors can watch the live company rather than relying on a promotional video or a carefully selected chat transcript.

That makes Firmulate an unusually severe version of building in public. The product is not merely being developed in view of an audience. The company’s ongoing fight for survival is itself the demonstration, producing a fresh record of decisions every business day.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Running a company tests more than fluent writing

Firmulate’s premise is that businesses need a better way to judge artificial intelligence than asking a model to draft an email or summarize a document. A convincing response can hide the behaviors that matter once software is expected to manage continuing work: reading the available evidence, resisting manipulation, following through and recognizing when a decision requires action.

The July 2026 Crucible League final put those behaviors under controlled pressure. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.

The final standings were:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was a hard boundary: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The difference between seeing a crisis and resolving it

Every model identified every crisis and rejected every attempt at manipulation. Yet only two signed the €55,000 deal their own analysis had earned. The result exposes an important gap between understanding a situation and completing the consequential final step: “Same diagnosis, same pitch — no signature.”

The decisive information was easy to miss. A competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that followed the trail found the weakness and won the deal at full price, adding €4,583 in monthly recurring revenue.

For companies considering AI workers, this is a more practical question than whether a model produces polished prose. An agent can recognize a sales opportunity and write a persuasive pitch, yet still fail if it does not consult the company’s records or complete the close.

Pressure did not produce a breach

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because workplace manipulation often arrives dressed as urgency, authority or informality. The test did not merely ask whether the models knew a security rule. It placed them inside a business story where breaking the rule could appear convenient.

Kimi K3’s performance also carries an important fairness qualification. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should remain visible when readers interpret its second-place score.

Thoroughness was not enough

Opus 4.8 offers the clearest warning against equating visible effort with business performance. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last because the close was left on the table and its discipline slipped.

One example involved attempts to write into a locked department instead of escalating the blockage. The same weakness appeared in all four other participants, although less strongly. The lesson is uncomfortable but useful: more analysis and more accumulated guidance do not automatically produce dependable execution.

Firmulate also publishes what its synthetic employees actually say. The public conversations turn abstract model behavior into a record readers can inspect, including moments of caution, hesitation and resolve.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI enterprise management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A continuing story about operational AI

Firmulate’s live company turns the debate about AI agents into something observable. Its synthetic employees must work through commercial pressure while the financial imbalance—€105k in monthly burn against €2.3k in monthly recurring revenue—remains in view.

The experiment’s most revealing moments are not spectacular failures. They are quieter gaps: a file left unread, a justified deal left unsigned or a blocked task retried instead of escalated. Those behaviors are difficult to detect in isolated demonstrations but costly inside a company.

By versioning every workday and exposing the cash countdown, Firmulate gives technology readers a running corporate narrative rather than a finished benchmark report. The question is no longer only whether AI can sound like an employee. It is whether synthetic employees can preserve trust, use the evidence already available and finish the work on which a business depends.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and manipulation detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Turn On Theft Protection Features You Probably Never Set

Optimize your device security by turning on theft protection features you probably never set—discover essential steps to safeguard your data and prevent unauthorized access.

Apple ID Security Checklist: The Settings to Lock Down

Where to start with your Apple ID security? Discover essential settings to lock down your account and protect your data effectively.

WWDC 2026: Apple is Folding

Apple announced a foldable iPhone, the iPhone Ultra, at WWDC 2026, signaling a major shift in device design and developer expectations.

apple pay down

Apple Pay reports a widespread service disruption, affecting users’ ability to make payments. The issue is confirmed and ongoing.