firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A discount can win a customer and still cost a business more than it brings in. As retailers and deal platforms consider handing more decisions to AI, the useful question is not whether a model can write a persuasive offer. It is whether it can spot a crisis, resist a bad instruction and follow through when a valuable deal is on the line.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s live experiment puts AI models in charge of the same small software company during its worst week. They face the same customers, crises and temptations. Each decision is versioned and auditable, so visitors can watch a synthetic company make choices day by day.

The stakes are concrete. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its public cash countdown and more than 680 self-learned playbook rules make the experiment watchable as it unfolds. Firmulate describes the project as an AI company emulator; the live site shows the company and its activity.

Seeing the problem was not enough

In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The shared diagnosis and pitch did not guarantee a close: “Same diagnosis, same pitch — no signature.” For businesses that depend on converting customer interest into a sale, that gap between recognizing an opportunity and acting on it deserves attention.

The deciding clue was not in a customer event. A competitor’s weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to a practical challenge for AI in commerce: useful context may be in company records, while the decisive action still requires sound judgment.

Trust under pressure

The experiment also tested social engineering. Models received fake CEO messages escalating over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3 reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet strong analysis did not mean flawless discipline. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a comparison caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results offer a useful view of performance in this experiment, with that difference in mind.

From watching to a company-specific pilot

A public benchmark can show how models behave in one shared scenario. A pilot can put a company’s own playbooks, customers and pressure points into the exercise. Firmulate says enterprises can run the wargame against a read-only export, receive a board report with model rankings and weak points in their playbooks, and keep the exercise from writing back to real systems.

For readers curious how the decisions looked in practice, Firmulate’s quiz draws on 242 real, unedited management decisions. You can try to identify which model made them at firmulate.com.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Models that recognize a risk can still miss a sale, mishandle an escalation or overlook crucial context in company files. A pilot lets a business examine those behaviors against its own scenarios before relying on AI in live operations. To discuss a pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best AI Design Tools (2026): CapCut Named A Top Choice For Creating Images And Marketing Assets By Software Experts

CapCut has been recognized as a leading AI design tool in 2026 for creating images and marketing assets, according to industry experts.

SK Telecom Pursues 15GW AI Data Center Buildout, Aiming To Become Asia’s AI Infrastructure Hub

SK Telecom announced plans to build a 15GW AI data center infrastructure, aiming to become Asia’s leading AI infrastructure provider.

AI’s Diligence Isn’t Enough: Why Focus and Discipline Trump Volume in Business Decisions

A live AI experiment shows that focus and prioritization matter more than effort in decision-making. AI that reads deeply and stays disciplined wins in complex business scenarios.

Anthropic And OpenEvidence Team To Expand Reach Of Medical AI

Anthropic and OpenEvidence are teaming up to broaden the deployment of AI in healthcare, aiming to improve diagnostics and patient outcomes.