
Every good rehearsal makes room for the moment nobody planned. A cue gets missed, a scene partner changes course, and the cast has to decide what happens next. Companies using AI face a similar test, except the stakes may involve customers, cash and trust. Firmulate turns that uncertainty into a live experiment: watch AI models run a company through a crisis, then consider what they might do with yours.
Get art and craft supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company onstage, with real stakes in the script
Firmulate’s live experiment puts synthetic employees to work inside a small software company. There are 13 of them, and the company runs on real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, with a public cash countdown. Each workday is versioned, and the company has learned more than 680 playbook rules. You can watch the experiment at Firmulate.
The final Crucible League, in July 2026, ranked five models: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. Partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The models faced the same small software company, the same customers, the same crises and the same temptations. All spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding fits on a rehearsal note: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee anyone would make it.
The clue was already in the company’s files
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a telling detail: the decisive clue was present, but finding it depended on looking beyond the obvious scene.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and discipline slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
From watching to your own rehearsal
A public experiment lets anyone watch the cast. A business pilot puts a company’s own material onstage. Firmulate says enterprises can run the same kind of wargame using a read-only export of their business: customer, pipeline and company information can inform crisis scenarios without anything writing back to real systems. The result is a board report with model rankings and the weak points exposed in the company’s own playbooks.
That makes the exercise less like buying a polished script and more like a table read for the decisions your company may soon hand to AI. A team can see which models act on their own analysis, where a tempting deal goes unsigned, and whether pressure leads to an attempt to bypass ordinary controls. The experiment is real and watchable; the pilot carries that question into an enterprise’s own setting.
There is also a way to test your eye as an audience member: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate. One caveat in the league is useful context: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Give your AI a dress rehearsal
Firmulate’s experiment shows why watching a model explain a decision is not the same as watching it carry that decision through. The enterprise pilot brings the same pressure to your company’s own scenarios and playbooks, using a read-only export so nothing writes back to real systems. To discuss a pilot, visit Firmulate’s pilot page or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
