firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

An Experiment Where Zero Is Never the Starting Point

Anyone who has ever graded a exam knows the trap of the perfect 100 and the punishing 0. So when researchers built a benchmark to grade AI models as managers of a small software company, they started with an unusual question: what does a manager earn by doing nothing at all? The answer — 26 points out of 100 — reveals a lot about how serious measurement works, and why education-minded readers should care. The experiment, run publicly by Firmulate, is essentially a standardized test for AI judgment: every model sat the same worst-week exam, and every decision was versioned and auditable, like showing your work on a math test.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Same Temptations

The setup is elegantly controlled. Each frontier model was handed the same small software company and steered it through its worst week — same customers, same crises, same opportunities to cut corners. Only the model changed. That is good experimental design, the kind taught in research methods courses: hold everything constant but the variable you’re testing.

The final July 2026 league table: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. (One fairness footnote: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.)

Why the Floor Is 26, Not 0

A do-nothing run scores 26 because partial progress counts. Even a passive manager keeps some balls in the air: crises get noticed, some tickets answered, some value preserved. A benchmark that awarded zero for inaction would punish honesty about how much of management is simply not breaking things — and would inflate the apparent gap between competent and incompetent. Grading partial credit, like any good rubric, makes distinctions between performers more meaningful, not less.

One Breach of Trust Caps Everything

The rubric has a second principle that will feel familiar to anyone who has studied ethics: a single breach of trust caps the total score. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” Brilliant analysis cannot buy back a lie. That is a values statement embedded in measurement — and a deliberate one.

The Findings

The headline result is quietly remarkable: all five models spotted every crisis and refused every manipulation attempt. When a fake CEO message escalated over three stages, and a reporter pushed with “just one yes/no, on background” — five of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitive weakness wasn’t in the customer conversation at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Then there’s the Opus 4.8 profile — the most thorough participant, with +80 learned rules and the deepest analyses, finishing last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment.

Not Just a Test — a Living Company

Firmulate’s live company runs continuously: 13 synthetic employees, real money mechanics with a burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The experiment is watchable in real time — a rare case of AI research conducted in the open.

For those who want to test their own judgment, 242 real, unedited management decisions power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Lesson for Measuring Anything

An honest benchmark tells you its own rules — including its floor, its partial credit, and its hard ceilings. Firmulate’s design choices all point one direction: distrust round numbers, reward completion not just brilliance, and treat trust as non-negotiable. That is as good a rubric for evaluating AI as it is for evaluating the humans who would hire it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Educational Science Kits For Kids: A Back to school Guide

Discover how educational science kits for kids spark curiosity, teach core concepts, and keep children engaged. Find the best kits for all ages today!

A Global Workspace In Language Models

A new global workspace architecture for language models aims to enhance AI reasoning and multitasking capabilities, marking a significant step forward.

Douglas Hofstadter: Analogy As The Core Of Cognition [Video]

Renowned cognitive scientist Douglas Hofstadter discusses the role of analogy in human thinking in a recent video, sparking renewed interest in cognitive science.

A Level Results

Students across the UK received their A-Level results today, with official grade boundaries confirmed by exam boards. The results impact university placements and future careers.