
Imagine buying a new kitchen gadget that claims to be the “best” but often fails just to turn on. In the world of AI, benchmarks reveal a similar story: even the most basic models can score surprisingly high, raising questions about what these numbers really mean for your business. As AI tools become more integrated into operations—handling customer service, forecasting, or support—it’s crucial to understand what these scores actually tell us about their reliability and trustworthiness.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What is an AI Benchmark, and Why Does a ‘Do-Nothing’ Model Score 26?
In a recent public experiment called the Crucible League, four advanced AI models faced the same challenging week of running a small software company. The goal? See if they could navigate crises, uphold honesty, and close deals. Surprisingly, even the most passive or minimal effort models scored around 26 points — not zero. This baseline score reflects that even a simple, do-nothing approach garners partial credit, simply by not failing outright.
This setup is essential because it shows that AI assessments aren’t simply about whether an AI produces text or answers questions. Instead, they measure how well AI systems perform in complex, real-world scenarios—where trust, thoroughness, and discipline matter just as much as accuracy.
business AI trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Counts and Trust Caps the Score
In the experiment, all four models identified every crisis and refused manipulation attempts—demonstrating they can recognize threats and maintain integrity. However, only two of them went further: they signed a €55,000 deal their own analysis had earned. The other two, despite diagnosing correctly, didn’t follow through and left the deal on the table.
Here’s the key: partial success—like recognizing a crisis—is rewarded, but a breach of trust caps the maximum score. If an AI attempts manipulation or breaks protocol, it’s disqualified from higher points, regardless of other strengths. This approach emphasizes honesty and discipline, qualities as vital as problem-solving in a business setting.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed About AI’s True Capabilities
The models were tested on their ability to read and understand critical internal documents, not just surface-level customer interactions. In fact, the decisive advantage went to the models that read two document references deep into the company’s files. Those models secured the deal at full price, adding over €4,500 in monthly recurring revenue to the business.
They also faced social engineering—fake CEO messages and reporter tricks—where every model refused to be manipulated. Kimi K3, for example, explained: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating cautious judgment that aligns with real-world trust standards.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business Managers
Many businesses are eager to adopt AI, but they often judge performance by flashy chat demos or impressive language. The truth is, for operational tasks—managing support, sales, or decision-making—what matters is whether the AI can finish what it starts, read critical documents, and stay honest under pressure.
In the Crucible League, the models ran a simulated company with 13 synthetic employees and real-money mechanics. Despite burning over €105,000 a month against a low €2,300 monthly revenue, the models demonstrated discipline and strategic decision-making—just like a seasoned manager.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Significance of a Real-World Benchmark
What sets Firmulate’s benchmark apart is transparency and honesty. A do-nothing baseline score of 26 shows that even minimal effort is partial progress, and crossing the line into trust violations caps the maximum score. This ensures that AI developers and users focus on meaningful attributes—trustworthiness, thoroughness, and discipline—not just language prowess.
For companies considering AI, the takeaway is clear: the true value lies not just in AI’s ability to generate convincing text but in its capacity to make sound decisions, read internal files, and remain honest under pressure. The current league table highlights that even the best models are still learning these qualities, and rigorous testing like this is essential before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
