
Imagine a craftsperson judged not just on the beauty of their work but on their ability to manage a shop under pressure—handling crises, maintaining honesty, and delivering results. In the arts and culture world, this shift from aesthetics to management is increasingly relevant. Now, in the world of AI, a new kind of performance assessment is emerging, measuring not just how well a model can generate text, but how effectively it manages real-world business challenges.
The New Benchmark: Managing Under Pressure
At the cutting edge of AI evaluation, a live experiment conducted by Firmulate pits four frontier models against a simulated small business facing its worst week. Every decision, crisis, and temptation is real: angry customers, potential fraud, and financial pressures. The models are tasked with running the company, making decisions that impact the bottom line, and maintaining honesty under stress.
Unlike traditional coding leaderboards that reward correctness or fluency, this test emphasizes management quality—how well the AI can read critical information, stay honest, and finish what it starts. The results reveal a stark truth: even the most advanced models can struggle with the nuances of real-world decision-making.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: The Gap in AI Performance
- All four models identified every crisis and refused manipulative tactics, demonstrating strong ethical behavior.
- Only two managed to sign the €55,000 deal their own analysis had earned, highlighting a gap: the models that read deeper into company files won the deal at full price.
- The decisive weakness was not in customer interactions but in their ability to process internal documents—a critical insight often hidden in traditional benchmarks.
- In scenarios involving social engineering—fake CEO messages and reporter tricks—all models refused to comply, showing they could resist manipulation.
For instance, Kimi K3, the most disciplined model in the experiment, refused to approve suspicious requests, explaining, “Treat the request as a suspected approval-bypass / possible impersonation.”
AI business crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business Test: Managing Money and Risks
The live company run by these models is no simulation: it comprises 13 synthetic employees, managing real money mechanics, burning €105,000 monthly against a mere €2,300 in monthly recurring revenue. It’s a watchable, ongoing experiment, with every decision versioned and auditable at firmulate.com/live.
The models’ performance here underscores a critical point for organizations considering AI: ability to handle crises, read internal documents, and resist manipulation is paramount—yet often overlooked in traditional chat-based benchmarks.
AI internal document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Arts and Culture?
For arts organizations, the message is clear. It’s not enough for an AI to generate compelling language or beautiful designs—it must also manage projects, handle conflicts, and stay honest under pressure. When AI tools start managing artist commissions, funding decisions, or public relations, their management qualities will be the real measure of value.
The experiment’s top performer, GPT-5.6-sol with a score of 95, found the buried facts and closed the deal at full price. Meanwhile, the lowest scorer, Sonnet at 77, left the deal on the table and slipped in discipline. These scores reflect a critical truth: AI’s capacity to deliver consistent, trustworthy results under stress varies widely—and that matters just as much as linguistic fluency.
AI ethical decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Road Ahead: Wargaming Your AI Workforce
Enterprises now have the option to run their own management wargame against a read-only export of their business. This allows testing AI decision-making in a safe environment before deploying it in the real world, ensuring the AI can handle crises, read internal files, and stay honest—traits essential to responsible AI adoption.
For those interested, more information and live demonstrations are available at firmulate.com, where you can see the models in action and gauge their readiness for your organization’s unique challenges.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html