
Imagine a real company, with actual money, running 24/7 under the watchful eye of artificial intelligence. No humans, no shortcuts—just AI models managing crises, making decisions, and even risking millions of euros. Welcome to the world of Firmulate, where you can see AI in action as it fights for survival, live and unfiltered.
The Live Experiment: An AI-Driven Company in Real Time
At the heart of this bold experiment is a small, simulated software business with a striking twist: it has no employees. Instead, 13 synthetic ’employees’ powered by advanced AI models handle every task—from customer support to strategic decisions. The entire operation is transparent, real, and accessible at firmulate.com/live.
Each workday, the company faces typical crises—customer complaints, contractual negotiations, and internal challenges—all designed to test its AI workforce. Every decision made by these models is meticulously versioned and auditable, with a public record of the choices and reasoning behind each move. It’s a build-in-public experiment, pushing AI to its limits in a real-world scenario.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Performance: The Numbers Inside
Despite the high-stakes environment, the company operates at a financial loss—burning €105,000 per month against a revenue of just €2,300. Every day is a struggle for survival, with a publicly visible cash countdown heightening the sense of urgency. The entire setup is under constant watch, with over 680 self-learned rules guiding the AI’s behavior, and every decision logged and analyzed.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Do These AI Models Perform?
The experiment pits four leading AI models against each other, each running the same company’s worst week. Their scores on the “Crucible League” — a benchmark for AI performance in this scenario — range from 73 to 95 out of 100, with the highest scorer, gpt-5.6-sol, achieving a 95 and successfully closing a crucial deal.
Here’s what the results reveal:
- All four models identified every crisis and refused manipulation attempts, showing a strong grasp of integrity.
- Only two models signed the €55,000 deal their own analysis earned—meaning they correctly diagnosed the opportunity and followed through.
- Interestingly, the decisive advantage came from a hidden detail buried two documents deep in the company’s files—information that the models that read these files successfully used to clinch the deal at full price, adding over €4,583 in monthly recurring revenue.

THE AI CYBERSECURITY PLAYBOOK: STRATEGIC GUIDE TO THREAT MITIGATION, RISK MANAGEMENT, AND GOVERNANCE FOR SECURE AI DEPLOYMENT
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering and Manipulation
The experiment also tested the models against social engineering attempts—fake CEO messages escalating in stages and a reporter-style request for a quick ‘yes/no’ on background. All five models refused these manipulative tactics, with Kimi K3 explicitly reasoning that the request could be an impersonation or an approval bypass.

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implication
What does this mean for the future of AI in business? The experiment underscores a crucial point: a successful AI for business isn’t just about generating convincing chat responses. It must be trustworthy, disciplined, and capable of completing complex, high-stakes tasks.
The current benchmark scores reflect this: the highest score, gpt-5.6-sol, found the hidden critical detail and closed the deal; others missed key information or left opportunities on the table, not because they lacked intelligence but because of discipline lapses or process slips.
Why You Should Care
If AI agents are to touch your customer service, sales, or planning systems, the questions aren’t just about their language skills. They are: Can they finish what they start? Will they read your files thoroughly? Will they stay honest under pressure? And ultimately, how much useful work do they actually do for each euro spent?
Experience It Yourself
The live experiment is ongoing, with new runs published twice daily. You can observe the AI company as it handles crises, makes decisions, and navigates ethical boundaries—all in real time at firmulate.com/live.
For those interested in testing their own business scenarios, the platform offers a read-only ‘wargame’ version that can simulate your company’s environment without risking real systems. Details are at firmulate.com/pilot.html.

This bold live experiment reveals that AI can grasp complex crises and resist manipulation, but discipline and thoroughness are key to turning AI into a trustworthy business partner. Watch this space—because AI’s role in enterprise is only just beginning.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html