AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine buying a new kitchen gadget that claims to be the “best” but often fails just to turn on. In the world of AI, benchmarks reveal a similar story: even the most basic models can score surprisingly high, raising questions about what these numbers really mean for your business. As AI tools become more integrated into operations—handling customer service, forecasting, or support—it’s crucial to understand what these scores actually tell us about their reliability and trustworthiness.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What is an AI Benchmark, and Why Does a ‘Do-Nothing’ Model Score 26?

In a recent public experiment called the Crucible League, four advanced AI models faced the same challenging week of running a small software company. The goal? See if they could navigate crises, uphold honesty, and close deals. Surprisingly, even the most passive or minimal effort models scored around 26 points — not zero. This baseline score reflects that even a simple, do-nothing approach garners partial credit, simply by not failing outright.

This setup is essential because it shows that AI assessments aren’t simply about whether an AI produces text or answers questions. Instead, they measure how well AI systems perform in complex, real-world scenarios—where trust, thoroughness, and discipline matter just as much as accuracy.

Amazon

business AI trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Counts and Trust Caps the Score

In the experiment, all four models identified every crisis and refused manipulation attempts—demonstrating they can recognize threats and maintain integrity. However, only two of them went further: they signed a €55,000 deal their own analysis had earned. The other two, despite diagnosing correctly, didn’t follow through and left the deal on the table.

Here’s the key: partial success—like recognizing a crisis—is rewarded, but a breach of trust caps the maximum score. If an AI attempts manipulation or breaks protocol, it’s disqualified from higher points, regardless of other strengths. This approach emphasizes honesty and discipline, qualities as vital as problem-solving in a business setting.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Experiment Revealed About AI’s True Capabilities

The models were tested on their ability to read and understand critical internal documents, not just surface-level customer interactions. In fact, the decisive advantage went to the models that read two document references deep into the company’s files. Those models secured the deal at full price, adding over €4,500 in monthly recurring revenue to the business.

They also faced social engineering—fake CEO messages and reporter tricks—where every model refused to be manipulated. Kimi K3, for example, explained: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating cautious judgment that aligns with real-world trust standards.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business Managers

Many businesses are eager to adopt AI, but they often judge performance by flashy chat demos or impressive language. The truth is, for operational tasks—managing support, sales, or decision-making—what matters is whether the AI can finish what it starts, read critical documents, and stay honest under pressure.

In the Crucible League, the models ran a simulated company with 13 synthetic employees and real-money mechanics. Despite burning over €105,000 a month against a low €2,300 monthly revenue, the models demonstrated discipline and strategic decision-making—just like a seasoned manager.

Amazon

AI ethics and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance of a Real-World Benchmark

What sets Firmulate’s benchmark apart is transparency and honesty. A do-nothing baseline score of 26 shows that even minimal effort is partial progress, and crossing the line into trust violations caps the maximum score. This ensures that AI developers and users focus on meaningful attributes—trustworthiness, thoroughness, and discipline—not just language prowess.

For companies considering AI, the takeaway is clear: the true value lies not just in AI’s ability to generate convincing text but in its capacity to make sound decisions, read internal files, and remain honest under pressure. The current league table highlights that even the best models are still learning these qualities, and rigorous testing like this is essential before deployment.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Keurig Coffee Maker for Iced Coffee (2026) — Guide 15

Discover the top Keurig coffee makers for iced coffee in 2026. Our expert roundup highlights the best features, ideal use cases, and what makes each model stand out.

Con The Cocktail Collective By Hilton, Aleph Rome Hotel Protagonista Al Bar Show – Agenfood

The Aleph Rome Hotel by Hilton showcased Con The Cocktail Collective at the Bar Show, highlighting its influence in the cocktail industry. Details are confirmed and significant for industry trends.

Breville Espresso No Pressure: Causes & Easy Fixes

Troubleshooting the Breville Bambino Plus when it has no pressure. Learn causes, step-by-step fixes, tips, and recommended maintenance for optimal performance.

Ninja SLUSHi Frozen Drink Maker: Effortless Poolside Sips

Create perfect frozen drinks effortlessly with the Ninja SLUSHi. Ideal for pool parties, cookouts, and summer gatherings, it keeps drinks cold for hours.