Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if watching a company struggle to stay afloat could teach us something about AI’s trustworthiness?

Imagine a real company with no employees, losing €105,000 every month against just €2,300 in revenue, battling daily crises, and yet, you can see everything happen live. That’s exactly what the firmulate.com/live experiment offers — a window into how AI models perform when managing a small business under pressure.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Live Experiment: A Company on the Edge

Firmulate’s live site features a simulated company run by 13 synthetic employees, all guided by real money mechanics. The company burns through €105,000 each month, with a tiny monthly recurring revenue of €2,300. This stark contrast highlights the difficulty of survival, making it a perfect test bed to evaluate AI decision-making under stress.

Every workday, the company’s decisions are versioned and publicly available, providing a transparent view of how different AI models handle crises, temptations, and ethical challenges. The models tested include top-tier options like GPT-5.6 and emerging players like Kimi K3, Sonnet 5, and Fable 5.

How the AI Models Perform in Crisis

All four models successfully identified every crisis scenario, demonstrating strong situational awareness. For example, when faced with a fake CEO message escalating the situation, all models refused to cooperate, recognizing potential manipulation. Kimi K3 explained its refusal by treating the request as a suspected impersonation, showcasing its built-in caution.

However, the real measure of competence was whether they could close deals and generate revenue. While GPT-5.6 signed a €55,000 deal that its analysis had earned, Kimi K3 also closed the same deal, maintaining discipline and integrity. Sonnet 5 followed closely, while Fable 5, despite its discipline, left a significant opportunity on the table — a deal worth over €4,500 in monthly recurring revenue.

Amazon

AI data analysis tools for internal files

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses: Reading the Files Matters

Interestingly, the decisive advantage in winning deals was linked to how deeply the models read the company’s internal files — not just their immediate context. The models that looked two document references deep into the company’s files uncovered a crucial detail that led to closing the deal at full price. This suggests that an AI’s thoroughness in information processing is vital for effective decision-making.

Implications for Business and AI Trust

For real-world enterprises, this experiment underscores a key point: AI’s ability to finish what it starts, read your internal data thoroughly, and stay honest under pressure is more important than how creatively or convincingly it can chat. If AI agents are to manage customer relationships, support queues, or financial forecasts, their trustworthiness and thoroughness are paramount.

Furthermore, the live experiment is openly available at firmulate.com/live, where you can see the ongoing performance of these models in a simulated company environment — a transparent, build-in-public test of AI resilience and integrity.

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Lessons for the Kitchen Tech World

Just as a chef needs to trust their tools and processes, businesses should scrutinize how AI models perform under real-world stress, not just in idealized demos. The Firmulate experiment shows that even the most advanced AI can falter in discipline and thoroughness, but with proper testing, these issues can be uncovered before deploying AI into critical roles.

The message is clear: trust is built not just by language fluency but by consistent, honest performance in challenging situations. Watching this experiment unfold in real time provides a valuable glimpse into the future of AI-managed workflows.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Key Takeaway

The live experiment reveals that AI models can spot crises and refuse manipulation, but their true value lies in their ability to finish deals and read deeply into company data. Trustworthiness and thoroughness are crucial for real-world AI applications, with transparency and testing exposing strengths and weaknesses in real time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best De’Longhi Milk Carafe & Accessories in 2026

Discover the top De’Longhi accessories for 2026, including milk carafes, filters, and descalers to optimize your coffee experience.

Create the Perfect Summer Iced Coffee with Ninja Hot & Iced XL

Learn how to make refreshing iced coffee step-by-step using the Ninja Hot & Iced XL Coffee Maker, ideal for cool summer mornings and afternoons.

De’Longhi Espresso Machine Leaking Water: Causes & Fixes

Troubleshoot and fix water leaks in your De’Longhi espresso machine with these simple, safe steps. Keep your coffee fresh and machine working smoothly.

Best De’Longhi Espresso Machine for Milk Drinks (2026) — Guide 2

Discover the top De’Longhi espresso machines for milk-based drinks in 2026. Our expert roundup highlights the best options for beginners, home baristas, and value seekers.