AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a world where an AI that does nothing still scores 26 out of 100 in a performance test. Why does doing less still make the grade? The answer unveils a deeper truth about trust, honesty, and reliability in AI systems — especially when they’re asked to manage real-world crises.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get art and craft supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The True Measure of an AI’s Worth

At first glance, it might seem strange that a baseline AI — one that simply does nothing — can earn a score of 26 out of 100. This isn’t a flaw or a mistake; it’s a deliberate part of the benchmark’s design, reflecting the often complex and nuanced reality of AI decision-making in business. Every model, including the so-called ‘do-nothing baseline,’ is graded not just on what it achieves, but on how it handles the smallest, most critical details.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Still Counts

The benchmark challenges AI models by placing them in a simulated, high-pressure scenario: running a small software company through its worst week. All models face the same crises — from customer complaints to internal manipulations — and are evaluated based on their responses. Interestingly, even the do-nothing baseline scores 26 points, demonstrating that inaction still earns some recognition. This reflects a core principle: partial progress and cautious refusal to act unethically are valued more than reckless attempts or empty responses.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Breaches Cap the Score

A critical rule in this experiment is that a single breach of trust caps the total score at 26, regardless of other good decisions. This illustrates a fundamental truth: trustworthiness in AI isn’t just about what it does well, but also about avoiding harmful or dishonest actions. Even if an AI correctly identifies crises, a single slip — such as signing a questionable deal without proper analysis — can negate its entire performance. This design choice underscores the importance of honesty and integrity when deploying AI in sensitive roles.

Amazon

AI ethics and trustworthiness training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment in Action

Each AI model was tasked with managing the same set of scenarios, involving real issues like customer crises and manipulation attempts. The models demonstrated impressive vigilance: they identified all crises and refused every attempt at manipulation, including sophisticated social engineering tactics like fake CEO messages and reporter tricks. Only two of the models went as far as signing a €55,000 deal, which their own analysis had earned. The others, despite diagnosing accurately and refusing manipulation, failed to close the deal.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses in the Models

The experiment revealed an interesting vulnerability: the decisive failures weren’t in the obvious crisis points, but rather in the internal documents of the simulated company. The models that read and understood these deeper references succeeded in securing full-price deals, earning significantly higher scores. Conversely, models that failed to scan these documents left money on the table, illustrating that success hinges on thorough, multi-layered understanding — not just surface-level responses.

Social Engineering and Ethical Vigilance

All models showcased resilience against social engineering: staged attempts by a fake CEO and a reporter were all refused across the board. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights a crucial aspect of AI reliability — its capacity to recognize and reject ethically dubious requests, even under pressure.

The Real-World Company Behind the Test

The experiment takes place within a simulated company populated with 13 synthetic employees, managing real money mechanics that burn €105k monthly against a tiny €2.3k monthly revenue. Every decision made by the models is logged, versioned, and observable at firmulate.com/live. This transparency allows companies to watch how AI performs under real stress and make informed decisions about deploying these systems.

The Lessons for Business and Arts

This benchmark isn’t just about AI performance; it’s about trust, honesty, and the importance of thoroughness in decision-making. For artists, craftspeople, and cultural institutions, it echoes a timeless truth: true value isn’t just in quick results or surface appearances but in deep understanding and integrity. Just as a flawed artwork can betray its creator’s intent, an AI that breaches trust can undermine entire projects or organizations.

The Takeaway

In an age where AI is increasingly woven into the fabric of our work and creativity, knowing how these systems perform under pressure is critical. The Firmulate benchmark demonstrates that even a do-nothing AI can score decently, but only a trustworthy, thorough model can truly succeed — especially when real money, reputation, and ethical standards are at stake. As this live experiment continues to unfold, it teaches us that trustworthiness is the true currency of AI in business.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can’t-Miss Chicago Design Moments: Fall 2026 – Design.newcity.com

New City highlights key Chicago design moments for Fall 2026, sparking increased interest in the city’s evolving design scene amid rising coverage.

What Great Teams Understand About Accessible Forms

Inevitably, understanding accessible forms unlocks inclusive digital experiences that engage all users—discover how to elevate your team’s design practices today.

Why Mouse Fatigue Reduction Matters More Than Specs Alone

Caring about mouse fatigue reduction is crucial for long-term comfort and health, so discover how the right design can make all the difference.

Why Color Contrast Matters More Than Most Designers Think

Of all design elements, color contrast significantly impacts accessibility and user engagement, and understanding its importance can transform your approach to design.