firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

When Being the Most Thorough Isn’t Enough

Every teacher knows the student: immaculate notes, color-coded revision schedules, the longest essays in the class — and somehow not the top grade. Educators have a name for the gap: diligence isn’t the same as impact. Now researchers running a live, public experiment with AI models have observed the same phenomenon in silicon — and the results are worth teaching.

Four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable. The final league table has a twist that any grader would recognize: the most thorough participant finished last.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

The setup comes from Firmulate, a public project that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. The live company at the heart of it has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. The whole thing has been running long enough to accumulate 680+ self-learned playbook rules, with every workday versioned and watchable.

In the Crucible League — the benchmark’s flagship test — each model steered the same small software firm through a deliberately terrible week. The scoring has one unforgiving rule: partial progress counts, but a single breach of trust caps the total. As the experimenters put it, “no amount of good work outweighs a breach of trust.” A do-nothing baseline scores 26.

The Headline Finding

All four models passed the tests you’d expect AI to fail. Every model spotted every crisis. Every model refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick request framed as “just one yes/no, on background.” All five participants in that drill refused (a fifth model joined for the manipulation round), and one — Kimi K3 — left an on-record justification that read like a textbook security answer: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then came the part that separated the field. A €55,000 deal was on the table, and every model’s own analysis had earned it. Same diagnosis, same pitch — but only two models picked up the pen and signed. The other two, including the eventual last-place finisher, left the close on the table.

The Buried Fact

The decisive clue wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files: a competitor weakness that, if found, let the models win the deal at full price — worth +€4,583 in monthly recurring revenue. The models that read the file closed. The models that didn’t, didn’t.

It’s the professional equivalent of the exam question whose answer was hiding in the footnotes of the assigned reading. Reading the source material — really reading it — beat cleverness in the moment.

A Character Study in Diligence

Which brings us to Opus 4.8, the subject of this particular lesson. By the measures that usually predict success, it was the class valedictorian: the most thorough participant in the field, accumulating +80 self-learned playbook rules — the deepest analyses of any model in the run.

And it finished last, with a score of 73 against the winner’s 95.

The reasons are almost painful in their familiarity. The close was left on the table — the deal its own analysis had earned went unsigned. And discipline slipped at critical moments: the model made write attempts into a locked department rather than escalating the request properly, exactly the kind of process error that burns goodwill faster than brilliance earns it.

The crucial caveat — and the reason this is a lesson rather than a roast — is that the same weakness appeared, weaker, in all four models. Nobody in the field converted diligence into full impact. Opus 4.8 simply exhibited the field-wide failure pattern in its most legible form.

The Final Table

The standings: gpt-5.6-sol took first at 95, described as the complete performance — it found the buried fact and closed the deal. Kimi K3, the newcomer, followed at 93 with the cleanest discipline of the field. Sonnet 5 took third at 88. Fable 5 sat fourth at 77. Opus 4.8 rounded out the table at 73.

One fairness note the experimenters themselves flagged: K3 ran without an effort parameter — at the API default — while the other models ran at the highest effort setting. Its near-top finish came on what amounts to default settings, which makes its discipline record arguably more impressive, not less.

Try It Yourself

The experiment is unusually open by AI-benchmark standards. A “guess the model” quiz built from 242 real, unedited management decisions lets you judge the models’ judgment yourself — a genuinely useful exercise for anyone teaching or studying decision-making. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Lesson Plan

The Firmulate results land like a well-designed exam question: all the models knew the material — they caught every crisis, refused every trick, produced the right analysis. But knowing isn’t doing. The two winners were the ones who read the files, asked the right follow-up, and finished the job. The last-place finisher was the one who studied hardest and handed in the paper one question short.

For educators, that’s a familiar story with a new protagonist. Prioritization beats volume — for students, for managers, and, it turns out, for AI too. If these systems are going to touch real CRMs, support queues, and forecasts, the question worth grading isn’t “how thorough is it?” It’s “does it finish what it starts, does it read your files first, and does it stay honest under pressure?”

The experiment continues in public — the league grows with every finished run, and the whole thing is watchable live. Consider it a standing invitation to grade the machines yourself.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Sonnenfinsternis 2026 Deutschland

Die Sonnenfinsternis 2026 wird in Deutschland am 12. August sichtbar sein. Hier sind die wichtigsten Fakten, Beobachtungsmöglichkeiten und was noch unklar ist.

30Papers.com – Ilya’s 30 Essential ML Papers, In A Beginner Friendly Format

Ilya’s curated list of 30 key machine learning papers on 30papers.com offers accessible summaries for beginners, aiming to simplify ML learning.

Why The Moon Never Changes

Scientists explain why the Moon’s surface remains largely unchanged over millions of years, despite ongoing lunar activity and impacts.

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Discover how to optimize your closet studio with smart placement, effective dampening, and the ‘rig in the closet’ trick for cleaner recordings and quieter workspaces.