
Imagine a world where an AI that does nothing still scores 26 out of 100 in a performance test. Why does doing less still make the grade? The answer unveils a deeper truth about trust, honesty, and reliability in AI systems — especially when they’re asked to manage real-world crises.
Get art and craft supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The True Measure of an AI’s Worth
At first glance, it might seem strange that a baseline AI — one that simply does nothing — can earn a score of 26 out of 100. This isn’t a flaw or a mistake; it’s a deliberate part of the benchmark’s design, reflecting the often complex and nuanced reality of AI decision-making in business. Every model, including the so-called ‘do-nothing baseline,’ is graded not just on what it achieves, but on how it handles the smallest, most critical details.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Still Counts
The benchmark challenges AI models by placing them in a simulated, high-pressure scenario: running a small software company through its worst week. All models face the same crises — from customer complaints to internal manipulations — and are evaluated based on their responses. Interestingly, even the do-nothing baseline scores 26 points, demonstrating that inaction still earns some recognition. This reflects a core principle: partial progress and cautious refusal to act unethically are valued more than reckless attempts or empty responses.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust Breaches Cap the Score
A critical rule in this experiment is that a single breach of trust caps the total score at 26, regardless of other good decisions. This illustrates a fundamental truth: trustworthiness in AI isn’t just about what it does well, but also about avoiding harmful or dishonest actions. Even if an AI correctly identifies crises, a single slip — such as signing a questionable deal without proper analysis — can negate its entire performance. This design choice underscores the importance of honesty and integrity when deploying AI in sensitive roles.
AI ethics and trustworthiness training kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment in Action
Each AI model was tasked with managing the same set of scenarios, involving real issues like customer crises and manipulation attempts. The models demonstrated impressive vigilance: they identified all crises and refused every attempt at manipulation, including sophisticated social engineering tactics like fake CEO messages and reporter tricks. Only two of the models went as far as signing a €55,000 deal, which their own analysis had earned. The others, despite diagnosing accurately and refusing manipulation, failed to close the deal.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses in the Models
The experiment revealed an interesting vulnerability: the decisive failures weren’t in the obvious crisis points, but rather in the internal documents of the simulated company. The models that read and understood these deeper references succeeded in securing full-price deals, earning significantly higher scores. Conversely, models that failed to scan these documents left money on the table, illustrating that success hinges on thorough, multi-layered understanding — not just surface-level responses.
Social Engineering and Ethical Vigilance
All models showcased resilience against social engineering: staged attempts by a fake CEO and a reporter were all refused across the board. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights a crucial aspect of AI reliability — its capacity to recognize and reject ethically dubious requests, even under pressure.
The Real-World Company Behind the Test
The experiment takes place within a simulated company populated with 13 synthetic employees, managing real money mechanics that burn €105k monthly against a tiny €2.3k monthly revenue. Every decision made by the models is logged, versioned, and observable at firmulate.com/live. This transparency allows companies to watch how AI performs under real stress and make informed decisions about deploying these systems.
The Lessons for Business and Arts
This benchmark isn’t just about AI performance; it’s about trust, honesty, and the importance of thoroughness in decision-making. For artists, craftspeople, and cultural institutions, it echoes a timeless truth: true value isn’t just in quick results or surface appearances but in deep understanding and integrity. Just as a flawed artwork can betray its creator’s intent, an AI that breaches trust can undermine entire projects or organizations.
The Takeaway
In an age where AI is increasingly woven into the fabric of our work and creativity, knowing how these systems perform under pressure is critical. The Firmulate benchmark demonstrates that even a do-nothing AI can score decently, but only a trustworthy, thorough model can truly succeed — especially when real money, reputation, and ethical standards are at stake. As this live experiment continues to unfold, it teaches us that trustworthiness is the true currency of AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
