AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a world where artificial intelligence doesn’t just chat about art or craft but actually runs a business — making critical decisions under pressure, just like a human manager. For arts and culture enthusiasts, this might sound like science fiction, but it’s now a real-world experiment that uncovers the true personalities of cutting-edge AI models. The question isn’t just whether they’re smart — it’s whether they’re honest, disciplined, and reliable enough to handle the messy, high-stakes realities of business management.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Live Business Simulation: A New Frontier for AI Testing

At the heart of this experiment lies a unique setup: four leading AI models each steer a small software company through its worst week — a scenario packed with crises, tempting manipulations, and tight deadlines. The company isn’t fictional; it’s a real operation, with real money and live customer interactions, running every business day on a platform called Firmulate. Each model faces the same challenges, from customer disputes to internal ethical dilemmas, and all decisions are recorded for transparency and analysis.

Measuring Management, Not Just Conversation

Unlike typical AI demos, this isn’t about chat quality or clever responses. It’s about how well the models manage crises and uphold integrity. Every decision is a test of discipline, reading comprehension, and ethical stance, with the ultimate goal of closing deals and maintaining trust. The results are revealing: all four models identified every crisis and refused every attempt at manipulation, such as fake CEO messages or background approvals. Yet, only two managed to close a crucial €55,000 deal based on their own analysis — the kind of outcome that counts in real business.

Amazon

AI business decision making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Factor: Reading Into the Company Files

The decisive difference lay in one subtle factor: the models that read two document references deep into the company’s internal files secured the deal at full price, adding over €4,500 in monthly recurring revenue. Conversely, models that skimmed the surface or failed to delve into these internal documents left the deal on the table, missing vital clues. This underscores an essential insight: in complex decision-making, thorough reading and comprehension of internal data can be the difference between success and failure.

Behavior Under Social Engineering Attacks

In a clever twist, the models faced staged social engineering — escalating fake CEO requests and a reporter’s subtle background question. All five models refused to participate in approval bypasses or impersonation, with Kimi K3 explicitly citing a concern for possible impersonation. This shows these models are not only capable of detecting threats but also of reasoning about their own safety and ethics under pressure.

Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company: A Testing Ground for AI’s Business Fitness

The company in question runs with 13 synthetic employees, managing real money mechanics, losing €105,000 each month against a modest €2,300 monthly recurring revenue. The environment is intense: a live cash countdown, a set of 680+ learned rules, and every workday’s decisions being versioned and scrutinized. Watching this experiment unfold online offers a rare glimpse into how these models behave in a realistic setting, far beyond canned demos or abstract benchmarks.

The Profile of Opus 4.8 and the Lessons Learned

Among the models, Opus 4.8 was the most meticulous — analyzing extensively and applying over 80 learned rules — yet it left the deal on the table and showed discipline slip-ups, such as writing issues that should have been escalated. Interestingly, all models shared a similar weakness, suggesting that even the most thorough can falter without disciplined oversight. The experiment reveals a critical point: performance isn’t solely about raw knowledge but about disciplined application.

Amazon

enterprise AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Broader Implications: Trust, Compliance, and Cost

This experiment is more than just a game; it’s a mirror for what AI can and cannot do in real-world management. As AI agents begin to touch vital systems like customer relationship management (CRM), support queues, or forecasting, the key questions aren’t about their conversational skill but about their reliability, honesty, and ability to complete tasks without shortcuts. The models’ scores reflect this: the top performer, gpt-5.6-sol, scored 95 out of 100, with the ability to find critical information and close deals, while others like Sonnet 5 scored lower, indicating more slips and missed opportunities.

Can You Guess the Model? Try It Yourself

Interested in seeing how these models behave firsthand? On firmulate.com/quiz.html, you can participate in a live quiz featuring 242 real, unedited management decisions. Test your intuition and see if you can identify which AI model made each decision based on its management personality — whether it’s terse, thorough, or cautious. This isn’t just a game; it’s a window into the future of AI-driven management.

Amazon

AI ethical decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: The Human Side of AI Management

As arts and culture enthusiasts, it’s tempting to imagine AI models as silent artisans or creative collaborators. But as this live experiment shows, the real challenge is whether these models can act as responsible stewards of business, reading carefully, staying honest, and managing under pressure. The personalities they display — from meticulous analysts to disciplined skeptics — shape the future of how AI will integrate into our workplaces and creative industries. The takeaway? Trust isn’t just about how well an AI can chat; it’s about whether it can finish what it starts, stay disciplined, and earn your confidence.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Accessibility for Cognitive Disabilities: Simplifying Complexity

Learning how to simplify digital content for cognitive disabilities reveals transformative strategies that can truly enhance accessibility for everyone.

Can AI Save Your Business? Four Models, One Company, Two Victors

Four AI models ran a real company through its worst week; only two completed the critical deal. Testing real decision-making exposes AI’s true business strength, vital for arts and culture sectors.

Screen Reader-Friendly Interfaces: Tips and Best Practices

Optimize your web design for screen readers with essential tips and best practices to create inclusive, accessible experiences for all users.

How Readable Type Shapes Stronger User Experiences

AIThis post was created with the assistance of artificial intelligence (AI).Readable type…