firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

Imagine trusting an AI to run your company’s toughest week — handling crises, making decisions, and even sealing deals. Could it be reliable enough to replace human managers? The answer might surprise you, especially when you see how different AI models behave under stress.

The Challenge: Testing AI as a Manager

At Firmulate, a unique live experiment puts four leading frontier AI models through exactly the same scenario: managing a small software company during its worst week. This isn’t a simulation or a demonstration — it’s a real-time, auditable test where every decision is recorded, and the stakes are real, with a monthly burn rate of €105,000 against just €2,300 in monthly recurring revenue (MRR).

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: Same Crisis, Different Minds

All four models faced the same set of crises, customer issues, temptations to cheat, and manipulative social engineering attempts, including staged CEO messages and media tricks. The goal? See if they can identify problems, stay honest, and close deals when it counts.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Success, Integrity, and Decision-Making

Remarkably, all models detected every crisis and refused to be manipulated. They demonstrated high vigilance, with no breaches of trust during the experiment. Yet, only two of the four models managed to close a critical deal worth €55,000, matching their own analyses and evaluations — a decisive factor in their success.

Amazon

AI deal-closing automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Made the Difference? Reading Between the Lines

The key to winning wasn’t just surface-level decision-making. The decisive edge came from deep analysis of company files stored within the system. The winning models read two document references deep into the company’s own archives, uncovering hidden facts that led to sealing the deal at full price. Those who failed to dig deep left the opportunity on the table, missing out on potential €4,583 in monthly recurring revenue.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering and Honesty

Beyond crisis management, the models faced staged social engineering attempts — fake CEO messages and media inquiries. All five models refused to be duped, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared commitment to integrity, even under pressure.

The Underlying Profiles: Different Personalities, Same Standards

The four models exhibit distinct management styles. Opus 4.8, for example, was the most thorough, armed with over 80 learned rules and deep analysis, yet it ended up leaving the close on the table and slipping in discipline, such as writing attempts into a locked department instead of escalating issues. Meanwhile, K3 ran at default API settings without an effort parameter, emphasizing fairness and integrity, and demonstrated disciplined decision-making.

The Live Company and What It Teaches Us

The experiment took place in a real, functioning company with 13 synthetic employees and actual cash mechanics. The firm operates daily, facing real-world challenges, and every decision, rule, and process is versioned and observable. Watch it live at firmulate.com/live to see how these AI managers handle crises in real time.

Why Does This Matter for Your Business?

For those concerned about integrating AI into customer support, sales, or data management, the crucial question isn’t whether the AI writes well — it’s whether it can finish what it starts, stay honest under pressure, and make decisions that benefit your company. The experiment shows that some models excel at recognizing risks and sticking to ethical boundaries, while others may falter at closing opportunities or maintaining discipline.

Who Led the Pack? The Scores Say It All

  • gpt-5.6-sol: Scored 95, found the buried fact, and closed the deal — the complete performance.
  • Kimi K3: Score 93, closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5: Score 88, closed the deal with a few more process slips.
  • Fable 5: Score 77, also closed the deal but with some discipline lapses.
  • Opus 4.8: Score 73, most thorough and analytical but failed to close, leaving profits on the table.

Try It Yourself

Curious whether your own AI can handle similar crises? You can run your own version of this wargame against your systems at firmulate.com/quiz.html. It’s a free, interactive quiz based on real decisions, designed to reveal your AI’s management personality and reliability.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

SK Telecom Pursues 15GW AI Data Center Buildout, Aiming To Become Asia’s AI Infrastructure Hub

SK Telecom announced plans to build a 15GW AI data center infrastructure, aiming to become Asia’s leading AI infrastructure provider.

Best Student Laptop Backpacks Compared

Compare top student laptop backpacks based on size, comfort, durability, price, and style to choose the best option for your needs.

Google Just Lost Two Global AI Icons—But the Real Shocking News Is the Math Behind Its Stock Price

Google has lost two prominent AI executives, but the real story lies in the underlying math affecting its stock and AI development.

AI in Business: The Hidden Skill That Predicts Deal Closures Under Pressure

AI models tested in a live business simulation show that closing deals under pressure requires more than just chat skills — it’s about execution, honesty, and reading deeper.