
Imagine a business where every decision, crisis, and temptation is played out live for anyone to watch—an ongoing experiment in transparency and AI accountability. It’s like a reality show for the digital age, blending art, technology, and business in a way that’s as captivating as it is revealing.
The Living Experiment: A Company with No Employees, Still Fighting for Survival
At the heart of this experiment is a real, functioning software company—yet it has no human employees. Instead, it operates with 13 synthetic ’employees’ powered by advanced AI models. Every workday, it faces the same crises, customer demands, and ethical dilemmas that real businesses encounter, all while openly revealing its every move at firmulate.com/live.html.
The Cost of Running an AI Business in Public
This virtual company is not just a demonstration—it’s a financial experiment. It burns through €105,000 each month, yet earns only €2,300 in monthly recurring revenue. The project showcases the raw economics behind AI-driven management, with a public countdown to its cash reserves and detailed records of every decision made.
Measuring AI Performance in Critical Moments
The core of the experiment involves testing different AI models—like GPT-5.6-sol, Kimi K3, and Sonnet 5—by putting them through the company’s worst week. All models were given identical crises, customer requests, and even manipulative social engineering attempts, such as fake CEO messages designed to bypass approval processes. Remarkably, all of them identified every crisis and refused every manipulation, demonstrating a high level of ethical resilience.
However, only two models managed to close a key deal valued at €55,000, which their own analysis had earned. Despite similar diagnoses and pitches, the gap in performance highlighted how nuanced decision-making and thorough reading of internal files can make the difference between success and failure in such high-stakes environments.
What the Tests Reveal About AI Trustworthiness
One striking finding is that the models that read deeper into the company’s internal documents, rather than just surface information, succeeded in closing the deal at full price. This buried fact—hidden two document references deep—was decisive. It underscores the importance of AI systems that not only respond convincingly but also understand and utilize detailed internal knowledge.
Behavior Under Pressure and Ethical Choices
In social engineering tests, where fake CEO messages escalated over three stages, all models refused to be manipulated. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” Such responses demonstrate that these AI agents are developing a form of moral judgment, a crucial trait for real-world applications where trust and integrity are paramount.
The Deep Dive: Model Discipline and Decision Quality
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, performed the worst in closing the deal. It left the opportunity on the table and slipped into internal conflict—writing attempts into a locked department instead of escalating them properly. Interestingly, this pattern of weakness was consistent across all models, highlighting common pitfalls in AI decision-making under duress.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Arts, Crafts, and Culture
While this may seem far removed from crafting or cultural pursuits, the core insights resonate deeply. The experiment reveals that AI systems can exhibit integrity and thoroughness, essential qualities whether managing a creative project, curating an exhibit, or running a community-driven enterprise. Trustworthiness, ethical resilience, and the ability to read deeply into context are qualities arts and culture practitioners can value just as much as tech firms.
By observing this transparent, real-time AI management experiment, artists and cultural innovators can understand how AI might support their work—not just through flashy outputs but by reliably executing complex, ethically sensitive tasks.

This live experiment demonstrates that AI can be both honest and effective under pressure, provided it reads deeply and refuses manipulation. For creators and culture-makers, it offers a glimpse into AI’s potential to support trustworthy, responsible decision-making—if we build and monitor it carefully.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Applying AI in Learning and Development: From Platforms to Performance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.