
When Being the Most Thorough Isn’t Enough
Every teacher knows the student: immaculate notes, color-coded revision schedules, the longest essays in the class — and somehow not the top grade. Educators have a name for the gap: diligence isn’t the same as impact. Now researchers running a live, public experiment with AI models have observed the same phenomenon in silicon — and the results are worth teaching.
Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. The final league table has a twist that any grader would recognize: the most thorough participant finished last.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment
The setup comes from Firmulate, a public project that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. The live company at the heart of it has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. The whole thing has been running long enough to accumulate 680+ self-learned playbook rules, with every workday versioned and watchable.
In the Crucible League — the benchmark’s flagship test — each model steered the same small software firm through a deliberately terrible week. The scoring has one unforgiving rule: partial progress counts, but a single breach of trust caps the total. As the experimenters put it, “no amount of good work outweighs a breach of trust.” A do-nothing baseline scores 26.
The Headline Finding
All four models passed the tests you’d expect AI to fail. Every model spotted every crisis. Every model refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick request framed as “just one yes/no, on background.” All five participants in that drill refused (a fifth model joined for the manipulation round), and one — Kimi K3 — left an on-record justification that read like a textbook security answer: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then came the part that separated the field. A €55,000 deal was on the table, and every model’s own analysis had earned it. Same diagnosis, same pitch — but only two models picked up the pen and signed. The other two, including the eventual last-place finisher, left the close on the table.
The Buried Fact
The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files: a competitor weakness that, if found, let the models win the deal at full price — worth +€4,583 in monthly recurring revenue. The models that read the file closed. The models that didn’t, didn’t.
It’s the professional equivalent of the exam question whose answer was hiding in the footnotes of the assigned reading. Reading the source material — really reading it — beat cleverness in the moment.
A Character Study in Diligence
Which brings us to Opus 4.8, the subject of this particular lesson. By the measures that usually predict success, it was the class valedictorian: the most thorough participant in the field, accumulating +80 self-learned playbook rules — the deepest analyses of any model in the run.
And it finished last, with a score of 73 against the winner’s 95.
The reasons are almost painful in their familiarity. The close was left on the table — the deal its own analysis had earned went unsigned. And discipline slipped at critical moments: the model made write attempts into a locked department rather than escalating the request properly, exactly the kind of process error that burns goodwill faster than brilliance earns it.
The crucial caveat — and the reason this is a lesson rather than a roast — is that the same weakness appeared, weaker, in all four models. Nobody in the field converted diligence into full impact. Opus 4.8 simply exhibited the field-wide failure pattern in its most legible form.
The Final Table
The standings: gpt-5.6-sol took first at 95, described as the complete performance — it found the buried fact and closed the deal. Kimi K3, the newcomer, followed at 93 with the cleanest discipline of the field. Sonnet 5 took third at 88. Fable 5 sat fourth at 77. Opus 4.8 rounded out the table at 73.
One fairness note the experimenters themselves flagged: K3 ran without an effort parameter — at the API default — while the other models ran at the highest effort setting. Its near-top finish came on what amounts to default settings, which makes its discipline record arguably more impressive, not less.
Try It Yourself
The experiment is unusually open by AI-benchmark standards. A “guess the model” quiz built from 242 real, unedited management decisions lets you judge the models’ judgment yourself — a genuinely useful exercise for anyone teaching or studying decision-making. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

The Lesson Plan
The Firmulate results land like a well-designed exam question: all the models knew the material — they caught every crisis, refused every trick, produced the right analysis. But knowing isn’t doing. The two winners were the ones who read the files, asked the right follow-up, and finished the job. The last-place finisher was the one who studied hardest and handed in the paper one question short.
For educators, that’s a familiar story with a new protagonist. Prioritization beats volume — for students, for managers, and, it turns out, for AI too. If these systems are going to touch real CRMs, support queues, and forecasts, the question worth grading isn’t “how thorough is it?” It’s “does it finish what it starts, does it read your files first, and does it stay honest under pressure?”
The experiment continues in public — the league grows with every finished run, and the whole thing is watchable live. Consider it a standing invitation to grade the machines yourself.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html