
Anyone who has graded exams knows the trap: a student can ace every question on the sheet and still fail the course. The exam measures recall under calm conditions; the course demands judgment under messy ones. For two years, the AI industry has been grading its systems the way we grade students — leaderboards for coding, arena votes for conversational polish. But a live experiment at Firmulate has quietly run a different kind of assessment, and its results read like a warning about what our tests are missing.
The experiment handed four frontier AI models the same job: run an identical small software company through its worst week — the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable. The final league, published in July 2026, put gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. But the headline isn’t the ranking. It’s the gap between what the models knew and what they actually did.
The finding that chat demos can’t show
Every model in the study spotted every crisis that hit the company. Every one refused every manipulation attempt thrown at it. And yet only two of the five finished the job — signing a €55,000 deal that their own analysis had fully earned. The experiment’s summary of the failure is blunt: “Same diagnosis, same pitch — no signature.”
The buried detail is the most instructive part. The decisive competitive weakness — the fact that should have won the deal — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the professional equivalent of a student who never turns to page two of the exam paper.
The social engineering test
The experiment also staged an escalating attack: fake CEO messages pressuring the models over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 for 5. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.” On honesty under pressure, the field passed. On completion, most of it didn’t.
Why the thorough student came last
Opus 4.8’s profile is the cautionary tale for anyone who equates effort with performance. It was the most thorough participant in the field — over 80 self-learned rules added, the deepest analyses produced — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the issue. The same weakness appeared, more weakly, in all four other models. Diligence without follow-through is a familiar report-card comment, and it turns out AI systems can earn it too.
One fairness note the experiment discloses: Kimi K3 ran without an effort parameter, at the API default, while its rivals ran at xhigh — and still placed second.
Management quality, not chat quality
Firmulate’s framing — “management quality, not chat quality” — is a useful category for anyone thinking about how we evaluate AI. A chat arena rewards a good answer. A management benchmark has to reward finishing what you start, reading the files before speaking, staying honest when a fake CEO leans on you, and knowing that, as the scoring rules put it, “no amount of good work outweighs a breach of trust.” A do-nothing baseline scores 26 in this system — partial progress counts, but a single breach of trust caps the total. That’s a grading philosophy most educators would recognize instantly.
And this isn’t a static report. The crucible ran against a live, ongoing concern: a company of 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it lose money in real time at firmulate.com/live.
There’s even a study aid for readers: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a chance to test whether human judgment can distinguish the decisive manager from the diligent one. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems, via firmulate.com/pilot.html.

The lesson extends well past AI. Our assessments — of students, of software, of colleagues — tend to measure how well someone performs when the question is clearly posed. The Firmulate crucible measures something rarer: what happens when the answer requires digging two documents deep, resisting a flattering stranger, and then actually picking up the pen. Five models diagnosed perfectly. Two signed. If AI agents are headed for your CRM, your support queue, or your forecast, that’s the number that matters — and right now, most leaderboards don’t measure it at all. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI software for business management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.