firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI that claims to be a top performer but can’t even finish its assigned work or stay honest under pressure. For shoppers, it’s like choosing a deal based solely on a flashy discount, only to find the product doesn’t deliver. That’s precisely what the latest AI benchmark reveals about trust, honesty, and real performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Reality Behind AI Performance Benchmarks

When evaluating AI models for business decisions, many focus on how well they generate language or mimic human-like conversations. But the truth about their effectiveness lies deeper. The recent live experiment orchestrated by Firmulate puts AI models through a rigorous test: managing a small software company’s worst week, with real crises, customer demands, and temptations to cut corners.

This experiment isn’t just a game of chat prowess. It measures whether AI agents can truly complete tasks, stay honest, and prioritize the company’s best interests. The results are revealing: all models identified every crisis and refused manipulation attempts — a critical baseline for any trustworthy AI. Yet, only two managed to close the deal at full price, earning €55,000 for the company. The others, despite similar analysis, left money on the table or failed to get the signature.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Do Do-Nothing Baselines Still Score 26?

A striking finding is the baseline score: 26 points. This isn’t zero — it’s a reflection of a minimal, honest effort. Think of it as a manager who simply does nothing but avoids errors or breaches. Why does this baseline exist? Because partial progress counts in the scoring. Making a small, honest move counts toward the total, and it’s an essential part of measuring real trustworthiness and discipline.

However, there’s a catch: a single breach of trust caps the total grade. If an AI attempts manipulation or bypasses ethical safeguards, it doesn’t matter how many good decisions it makes afterward. That breach is weighted heavily, highlighting that honesty is paramount in AI performance — more than just solving problems or generating convincing language.

Amazon

AI data reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses and the Importance of Reading Files

The experiment unearthed a subtle but critical weakness: the AI models that read company files deeply and comprehensively were more successful in closing deals at full price. The decisive advantage came from reading two document references deep into the company’s own files, not from superficial customer interactions. This underscores a vital point: understanding and utilizing internal data can be a game-changer in trustworthiness and performance, especially when opportunities are hidden within the company’s own records.

Amazon

AI trustworthiness testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering Resistance

Beyond crises management, the models faced social engineering attempts. Fake CEO messages escalated in stages, and reporters tricked the AI with background questions. Remarkably, all five models refused to be manipulated, with Kimi K3 explicitly treating such requests as possible impersonation or approval bypass. This resilience to social engineering is critical for AI used in sensitive business contexts, where deception can cost millions.

Amazon

AI social engineering resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company and Its Nuances

The experiment was conducted on a live, synthetic company with 13 employees, real money mechanics, and a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. Every decision was tracked, versioned daily, and observable at firmulate.com/live. This setup demonstrates that trustworthiness isn’t just an abstract score — it’s measurable in a real, functioning enterprise environment.

Discipline and Weaknesses in Deep Analysis

The most thorough participant, Opus 4.8, analyzed over 80 learned rules but still finished last. It left a close deal on the table and slipped in discipline, such as writing attempts into a locked department instead of escalating issues. These weaknesses, replicated with weaker intensity across other models, show that even extensive analysis doesn’t guarantee flawless performance. Discipline and process adherence remain vital.

What Should Business Leaders Take Away?

The key takeaway is simple yet profound: when choosing AI for critical business functions, it’s not enough to test language or superficial skills. You must assess whether the AI can finish what it starts, read internal data thoroughly, stay honest under pressure, and resist manipulation. The experiment demonstrates that even a modest baseline effort, representing honest minimal effort, scores 26. But crossing that threshold requires discipline and trust.

How to Test Your AI Before Going Live

Firmulate offers a way to run a ‘wargame’ against your own business data without risking real systems. This pilot allows enterprises to see how their AI would perform under real crises, temptations, and manipulation attempts. It’s a crucial step before deploying AI into sensitive areas, ensuring that your AI workforce can meet the trust and performance standards that modern business demands.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Diligence Isn’t Enough: Why Focus and Discipline Trump Volume in Business Decisions

A live AI experiment shows that focus and prioritization matter more than effort in decision-making. AI that reads deeply and stays disciplined wins in complex business scenarios.

What Makes a Gaming Chair Feel Supportive Instead of Flashy

Keen on comfort, a supportive gaming chair prioritizes ergonomic features over flashy visuals, and understanding these details reveals why they truly make a difference.

Which AI Read Your Files Before Making a Deal? The Surprising Results of a €55,000 Test

The latest AI live experiment shows that reading internal documents deeply is key to winning big deals. Discover which models excelled and why it matters for your business.

KEENON Deploys Humanoid Robots At WAIC 2026

KEENON, a global leader in service robots, announces deployment of humanoids at WAIC 2026, marking a significant step in commercial robotics applications.