AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would you trust an AI with the dinner rush?

Imagine an appliance company facing a supplier delay, a wave of customer cancellations and a rival’s price cut all in the same week. An AI assistant might diagnose the problems perfectly. But would it follow through on the deal that could help the business—or hold its nerve when someone tries to trick it into breaking the rules? That’s the question behind Firmulate, a live experiment in what happens when AI models are asked to run a company.

Same company, same worst week

For its final Crucible League in July 2026, Firmulate put frontier AI models in charge of the same small software company and its worst week: identical customers, crises and temptations. Each decision was versioned and auditable. The point was to observe management under pressure, rather than judge how polished a model sounds in a chat.

Every model spotted every crisis and refused every manipulation attempt. Yet just two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” In the final standings, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The principle was blunt: “no amount of good work outweighs a breach of trust.”

The detail hidden in the paperwork

The deal hinged on a competitor weakness buried two document references deep in the company’s own files—not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a reminder familiar to anyone managing a kitchen supply chain: the decisive clue may be in a contract or account note, not in the latest alert.

The experiment also tested social engineering. Fake CEO messages escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still has to close

Opus 4.8 was the most thorough participant, learning more than 80 rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There’s a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions. Readers can try to guess which model made each choice.

From watching to a company’s own pilot

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. More than 680 self-learned playbook rules and every workday are versioned. The live experiment is watchable at firmulate.com, and the quiz is at Firmulate.

For businesses considering AI in customer support, sales or forecasting, the next step is to test decisions against their own operating context. Firmulate says enterprises can run the wargame using a read-only export of their business. The pilot does not write back to real systems. That offers a way to examine how models handle company-specific crises and where existing playbooks may leave gaps before agents get access to live workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

A model can identify a problem and still fail to act on its own advice. Firmulate’s experiment makes that gap visible—and offers enterprises a way to examine it against their own business data in a read-only pilot. Explore the pilot at firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Breville Portafilter Accessories & Parts in 2026

Discover the top Breville portafilter accessories parts in 2026. Find the best options for performance, value, and compatibility to upgrade your espresso game.

How Smart Kettles Improve Pour-Over Consistency

A smart kettle can transform your pour-over coffee experience, ensuring precision in temperature and timing—discover how it elevates your brew.

How Ice Makers Became Part of Premium Kitchen Culture

Kitchen ice makers have revolutionized premium culinary spaces, blending style with function—discover the unexpected ways they elevate your entertaining experience.

Is the De’Longhi La Specialista Worth It? Honest Review

An honest review of the De’Longhi La Specialista espresso machine, exploring its features, pros, cons, and who it’s best for in our detailed roundup.