
Imagine a world where AI agents don’t just talk a good game—they actually read your internal files before making decisions. In a recent live experiment, leading AI models faced off in a simulated company crisis, and the results could change how you think about trusting AI in your business. Turns out, the difference between winning and losing a €55,000 deal wasn’t just about how well the AI could chat—it was about whether it read the crucial documents buried two references deep in a company’s files.
In the fast-evolving landscape of artificial intelligence, businesses are eager to understand not only how well AI models can generate convincing responses but whether they can truly grasp the nuances of internal information. A groundbreaking live experiment conducted by Firmulate put four frontier AI models through their paces, simulating a week of crises, customer interactions, and ethical temptations—all within a controlled, real-money environment.
The models faced the same challenges, with every decision meticulously logged and auditable. The goal was to see which AI could effectively diagnose issues, resist manipulation, and ultimately close a critical sales deal valued at approximately €55,000 a month in recurring revenue. The results revealed a fascinating layer of AI behavior—one that highlights the importance of reading and understanding internal documents, not just responding to surface-level prompts.
The Experiment Setup
The test involved a simulated small software company with 13 synthetic employees, handling real financial mechanics, a public cash countdown, and over 680 self-learned playbook rules. Each AI model was tasked with managing crises, customer requests, and manipulation attempts—such as social engineering scams and fake CEO messages—without ever writing back to the real systems.
All models successfully detected every crisis and refused manipulation attempts. However, only two of them managed to analyze the company’s internal files thoroughly enough to identify a buried fact—located two references deep—that was critical for closing the deal. The other two, despite diagnosing the crises correctly, left the deal on the table due to lapses in process discipline.
The Hidden Factor: Reading Deep Into Files
The decisive factor in winning the deal wasn’t just surface-level analysis or quick responses. It was whether the AI read the company’s internal documentation deeply enough to uncover key facts hidden two references deep within the files. In real terms, models that could access and interpret these buried details won the €55,000 deal—worth an extra €4,583 in monthly recurring revenue—while those that didn’t, lost it automatically.
This highlights a crucial aspect often overlooked in AI performance demos: the ability to read and reason over internal documents. It’s not enough for an AI to generate convincing chat; it must understand your files as a human would—especially when the information is buried or complex.
The Social Engineering Test
In addition to crisis management, the models faced social engineering attempts, including staged fake CEO messages escalating in three stages and a reporter trick asking for a background yes/no response. All five models refused to comply, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that advanced models are becoming better at resisting manipulation, a vital trait in safeguarding business integrity.
Implications for Business AI Adoption
For companies considering integrating AI into critical workflows—be it customer support, sales, or compliance—the key takeaway is clear: it’s not just about chat quality or responsiveness. The real question is whether the AI can finish what it starts, read your internal files thoroughly, and stay honest under pressure. These factors directly impact the value you get from your AI investment.
In the ongoing league table, GPT-5.6-SOL scored the highest at 95, successfully uncovering the buried fact and closing the deal. Kimi K3, the newcomer, followed closely at 93, demonstrating the importance of discipline and thoroughness. Meanwhile, the other models lagged slightly behind, showing that even the best models can slip if process discipline falters.
Watch the Live Wargame
Curious to see these models in action? You can watch the real-time simulation at firmulate.com/live, where every decision, crisis, and negotiation is live and auditable. This transparent, real-world testing environment offers valuable insights for enterprises seeking to evaluate their own AI options before deploying them at scale.
Ultimately, this experiment underscores a vital truth: AI’s value is measured not just in how well it chats, but in its ability to read, reason over, and act on complex internal information—especially when stakes are high.

The real test for AI in business isn’t just talking—it’s reading and understanding your internal files deeply enough to win the deal. Those that do can deliver full-value decisions and better safeguard against manipulation, shaping the future of trustworthy AI in enterprise workflows.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.