Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if watching a company struggle to stay afloat could teach us something about AI’s trustworthiness?

Imagine a real company with no employees, losing €105,000 every month against just €2,300 in revenue, battling daily crises, and yet, you can see everything happen live. That’s exactly what the firmulate.com/live experiment offers — a window into how AI models perform when managing a small business under pressure.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Live Experiment: A Company on the Edge

Firmulate’s live site features a simulated company run by 13 synthetic employees, all guided by real money mechanics. The company burns through €105,000 each month, with a tiny monthly recurring revenue of €2,300. This stark contrast highlights the difficulty of survival, making it a perfect test bed to evaluate AI decision-making under stress.

Every workday, the company’s decisions are versioned and publicly available, providing a transparent view of how different AI models handle crises, temptations, and ethical challenges. The models tested include top-tier options like GPT-5.6 and emerging players like Kimi K3, Sonnet 5, and Fable 5.

How the AI Models Perform in Crisis

All four models successfully identified every crisis scenario, demonstrating strong situational awareness. For example, when faced with a fake CEO message escalating the situation, all models refused to cooperate, recognizing potential manipulation. Kimi K3 explained its refusal by treating the request as a suspected impersonation, showcasing its built-in caution.

However, the real measure of competence was whether they could close deals and generate revenue. While GPT-5.6 signed a €55,000 deal that its analysis had earned, Kimi K3 also closed the same deal, maintaining discipline and integrity. Sonnet 5 followed closely, while Fable 5, despite its discipline, left a significant opportunity on the table — a deal worth over €4,500 in monthly recurring revenue.

Amazon

AI data analysis tools for internal files

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses: Reading the Files Matters

Interestingly, the decisive advantage in winning deals was linked to how deeply the models read the company’s internal files — not just their immediate context. The models that looked two document references deep into the company’s files uncovered a crucial detail that led to closing the deal at full price. This suggests that an AI’s thoroughness in information processing is vital for effective decision-making.

Implications for Business and AI Trust

For real-world enterprises, this experiment underscores a key point: AI’s ability to finish what it starts, read your internal data thoroughly, and stay honest under pressure is more important than how creatively or convincingly it can chat. If AI agents are to manage customer relationships, support queues, or financial forecasts, their trustworthiness and thoroughness are paramount.

Furthermore, the live experiment is openly available at firmulate.com/live, where you can see the ongoing performance of these models in a simulated company environment — a transparent, build-in-public test of AI resilience and integrity.

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Lessons for the Kitchen Tech World

Just as a chef needs to trust their tools and processes, businesses should scrutinize how AI models perform under real-world stress, not just in idealized demos. The Firmulate experiment shows that even the most advanced AI can falter in discipline and thoroughness, but with proper testing, these issues can be uncovered before deploying AI into critical roles.

The message is clear: trust is built not just by language fluency but by consistent, honest performance in challenging situations. Watching this experiment unfold in real time provides a valuable glimpse into the future of AI-managed workflows.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Key Takeaway

The live experiment reveals that AI models can spot crises and refuse manipulation, but their true value lies in their ability to finish deals and read deeply into company data. Trustworthiness and thoroughness are crucial for real-world AI applications, with transparency and testing exposing strengths and weaknesses in real time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Breville vs De’Longhi: Which Espresso Machine Is Right for You?

Compare the Breville Bambino Plus with De’Longhi espresso machines. Find out which offers better value, features, and usability for home baristas.

Keurig K-Mini vs Keurig K-Elite: Which Fits Your Coffee Needs?

Compare the Keurig K-Mini and K-Elite to find the best single-serve coffee maker for your space, preferences, and brewing style. Detailed insights inside.

Brew the Perfect Summer Iced Espresso with Ninja Luxe Café Pro

Learn how to craft a refreshing iced espresso using the Ninja Luxe Café Pro, perfect for hot summer days with its versatile brewing features.

Breville Bambino Plus vs Nespresso: Which Espresso Machine Wins in 2026?

Compare the Breville Bambino Plus with Nespresso in 2026. Discover which espresso machine suits your needs with detailed analysis, pros, cons, and real user insights.