
Imagine if your favorite kitchen appliance could not only follow your recipes but also read and understand every note in your cooking files before suggesting a fix. Now, consider how much this level of thoroughness could mean for a business. Recent experiments show that AI models capable of analyzing files two references deep—essentially reading beyond the surface—are the ones closing the big deals. This isn’t about chatty AI; it’s about getting real work done by really understanding your data.
The Power of Deep Reading in AI Decision-Making
In a groundbreaking live experiment, four advanced AI models were tasked with running a small software company through its worst week—facing the same crises, same customers, same temptations. The goal? To see which AI could best diagnose issues, avoid manipulation, and ultimately close a €55,000 deal based on their recommendations. The results were telling: all four models identified every crisis and refused every manipulation attempt. Yet, only two managed to sign the deal, earning full recognition for their analysis.
As an affiliate, we earn on qualifying purchases.
What Made the Difference? How Deeply the AI Reads
The decisive factor wasn’t just how well the AI understood the surface data or responded to questions; it was whether it could read and understand information buried two document references deep within the company’s files. The models that achieved this depth of reading won the deal, boosting their revenue potential by over €4,583 monthly recurring revenue (MRR). In contrast, the models that missed this buried fact failed to sign the agreement, despite making the same diagnosis and pitch.
As an affiliate, we earn on qualifying purchases.
The Experiment’s Findings and Implications
- All models detected crises and refused manipulative tricks: They demonstrated a baseline ability to maintain integrity under pressure.
- Only two read deeply enough to find critical buried facts: The same key data that won the deal was hidden two document references down in the company’s files.
- Performance varied based on how thoroughly the AI read: The most comprehensive model, Opus 4.8, with over 80 learned rules and deep analysis, placed last—indicating that even thorough models can slip if discipline wanes.
- Fairness considerations: The Kimi K3 model ran without an effort parameter, making its performance directly comparable to others that ran at high effort levels.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Technology
For companies integrating AI into decision-making—whether in customer management, operations, or strategic planning—the key question isn’t just if the AI can generate good chat responses. It’s whether the AI can finish what it starts, understand the nuances buried in your files, stay honest under pressure, and do so reliably and consistently. The live experiment from Firmulate proves that AI’s true value lies in its ability to read and understand complex, multi-reference data, not just surface-level interactions.
As an affiliate, we earn on qualifying purchases.
The Real-World Application: Wargaming Your AI Workforce
Businesses can now simulate their own worst weeks using Firmulate’s AI company emulator—testing how different AI models handle crises, manipulations, and complex decision trees before deployment. This approach ensures that when AI interacts with real systems—support queues, CRM, forecasting—it does so with a proven capacity to deliver reliable, honest, and comprehensive work. The experiment underscores that AI’s decision quality hinges on its ability to engage with your data at multiple reference levels, not merely respond to prompts.

Deep reading capability in AI models is crucial for closing high-stakes deals and maintaining integrity under pressure. Firms that test their AI’s ability to understand complex documents before deployment will better ensure honest, reliable performance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html