
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would you trust an AI with the dinner rush?
Imagine an appliance company facing a supplier delay, a wave of customer cancellations and a rival’s price cut all in the same week. An AI assistant might diagnose the problems perfectly. But would it follow through on the deal that could help the business—or hold its nerve when someone tries to trick it into breaking the rules? That’s the question behind Firmulate, a live experiment in what happens when AI models are asked to run a company.
Same company, same worst week
For its final Crucible League in July 2026, Firmulate put frontier AI models in charge of the same small software company and its worst week: identical customers, crises and temptations. Each decision was versioned and auditable. The point was to observe management under pressure, rather than judge how polished a model sounds in a chat.
Every model spotted every crisis and refused every manipulation attempt. Yet just two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” In the final standings, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The principle was blunt: “no amount of good work outweighs a breach of trust.”
The detail hidden in the paperwork
The deal hinged on a competitor weakness buried two document references deep in the company’s own files—not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a reminder familiar to anyone managing a kitchen supply chain: the decisive clue may be in a contract or account note, not in the latest alert.
The experiment also tested social engineering. Fake CEO messages escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still has to close
Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There’s a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions. Readers can try to guess which model made each choice.
From watching to a company’s own pilot
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. More than 680 self-learned playbook rules and every workday are versioned. The live experiment is watchable at firmulate.com, and the quiz is at Firmulate.
For businesses considering AI in customer support, sales or forecasting, the next step is to test decisions against their own operating context. Firmulate says enterprises can run the wargame using a read-only export of their business. The pilot does not write back to real systems. That offers a way to examine how models handle company-specific crises and where existing playbooks may leave gaps before agents get access to live workflows.

Put your playbooks to the test
A model can identify a problem and still fail to act on its own advice. Firmulate’s experiment makes that gap visible—and offers enterprises a way to examine it against their own business data in a read-only pilot. Explore the pilot at firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
