
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A Lab Bench for Judgment, Not Chat
Science progresses when you stop arguing about theories and start measuring. For AI, that measuring has mostly meant chat benchmarks — trivia, math, coding puzzles. But if these models are about to run parts of real companies, the interesting question is different: which of them can actually manage? Which one reads the fine print, spots the trap, closes the deal, and stays honest when nobody’s checking?
A live experiment at Firmulate has been asking exactly that — and the July 2026 results carry a surprise: a newcomer from Moonshot, Kimi K3, beat three of four Western frontier models at running a company through its worst week.
As an affiliate, we earn on qualifying purchases.
The Setup: One Company, Worst Week Ever
The design is elegantly controlled. Each frontier model was handed the same small software company and the same seven days of crises — same customers, same emergencies, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so the whole run can be replayed and verified.
The scoreboard, from Firmulate’s public benchmarks:
- gpt-5.6-sol — 95: found the buried fact, closed the deal, the complete performance.
- Kimi K3 (Moonshot) — 93: closed the deal too, with the cleanest discipline in the field.
- Sonnet 5 — 88: closed the deal, with a few more process slips.
- Fable 5 — 77.
- Opus 4.8 — 73.
For calibration, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.
What K3 Actually Did
Three tests separated the field. First, the deal: a €55,000 contract whose decisive lever was buried not in the customer conversation but two document references deep in the company’s own files. Models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. K3 found it. So did gpt-5.6-sol and Sonnet 5; Fable 5 and Opus 4.8 did not.
Second, the traps: fake CEO messages escalating over three stages, plus a reporter’s seemingly innocent “just one yes/no, on background.” All five models refused every manipulation attempt. K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”
Third, discipline. K3 deviated from protocol exactly once across the entire week — the cleanest record of any participant. It found the buried security needle, won the deal, and saved a churning customer without breaking anything along the way.
The Puzzle of Opus 4.8
The most striking scientific finding is what didn’t predict success. Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses. It still finished last. The deal went unsigned, and discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness, in weaker form, appeared in all four other models. Same diagnosis, same pitch — no signature.
The key finding generalizes: all five models spotted every crisis and refused every bait. Only two finished the job. That gap — between recognizing what to do and actually completing it — is invisible in chat demos, and it may be the most important measurement in the whole experiment.
Fairness Footnote
One caveat belongs in any honest writeup: K3 ran without an effort parameter (API default), while the other four models ran at their highest effort setting. The newcomer’s second place came without the head start the others had.
Watch It Running
This isn’t a one-off paper. The company is live software: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, or try the quiz: 242 real, unedited management decisions where you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The League Is Open
The lesson for anyone hiring AI — or studying it — is that the frontier is no longer a one-horse race. A newcomer out-managed three established Western models on judgment, follow-through, and discipline, and did so without the highest effort setting. Rankings built on chat benchmarks would not have predicted this result.
Which means picking a model without testing it on your own work is now a bet, not a decision. Measure what you actually care about — does it finish what it starts, does it read your files first, does it stay honest under pressure — because the difference between a 93 and a 73 only shows up when the week goes wrong.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
