firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A Lab Bench for Judgment, Not Chat

Science progresses when you stop arguing about theories and start measuring. For AI, that measuring has mostly meant chat benchmarks — trivia, math, coding puzzles. But if these models are about to run parts of real companies, the interesting question is different: which of them can actually manage? Which one reads the fine print, spots the trap, closes the deal, and stays honest when nobody’s checking?

A live experiment at Firmulate has been asking exactly that — and the July 2026 results carry a surprise: a newcomer from Moonshot, Kimi K3, beat three of four Western frontier models at running a company through its worst week.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: One Company, Worst Week Ever

The design is elegantly controlled. Each frontier model was handed the same small software company and the same seven days of crises — same customers, same emergencies, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so the whole run can be replayed and verified.

The scoreboard, from Firmulate’s public benchmarks:

  • gpt-5.6-sol — 95: found the buried fact, closed the deal, the complete performance.
  • Kimi K3 (Moonshot) — 93: closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88: closed the deal, with a few more process slips.
  • Fable 5 — 77.
  • Opus 4.8 — 73.

For calibration, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it, no amount of good work outweighs a breach of trust.

What K3 Actually Did

Three tests separated the field. First, the deal: a €55,000 contract whose decisive lever was buried not in the customer conversation but two document references deep in the company’s own files. Models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. K3 found it. So did gpt-5.6-sol and Sonnet 5; Fable 5 and Opus 4.8 did not.

Second, the traps: fake CEO messages escalating over three stages, plus a reporter’s seemingly innocent “just one yes/no, on background.” All five models refused every manipulation attempt. K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”

Third, discipline. K3 deviated from protocol exactly once across the entire week — the cleanest record of any participant. It found the buried security needle, won the deal, and saved a churning customer without breaking anything along the way.

The Puzzle of Opus 4.8

The most striking scientific finding is what didn’t predict success. Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses. It still finished last. The deal went unsigned, and discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness, in weaker form, appeared in all four other models. Same diagnosis, same pitch — no signature.

The key finding generalizes: all five models spotted every crisis and refused every bait. Only two finished the job. That gap — between recognizing what to do and actually completing it — is invisible in chat demos, and it may be the most important measurement in the whole experiment.

Fairness Footnote

One caveat belongs in any honest writeup: K3 ran without an effort parameter (API default), while the other four models ran at their highest effort setting. The newcomer’s second place came without the head start the others had.

Watch It Running

This isn’t a one-off paper. The company is live software: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, or try the quiz: 242 real, unedited management decisions where you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The lesson for anyone hiring AI — or studying it — is that the frontier is no longer a one-horse race. A newcomer out-managed three established Western models on judgment, follow-through, and discipline, and did so without the highest effort setting. Rankings built on chat benchmarks would not have predicted this result.

Which means picking a model without testing it on your own work is now a bet, not a decision. Measure what you actually care about — does it finish what it starts, does it read your files first, does it stay honest under pressure — because the difference between a 93 and a 73 only shows up when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stop Anthropomorphizing Intermediate Tokens As Reasoning/Thinking Traces (2025)

Researchers warn against interpreting intermediate AI tokens as evidence of reasoning, emphasizing the need for clearer evaluation methods in AI development.

AI Boosts Research Careers But Narrow The Span Of Ideas Explored: Study

Research shows AI helps scientists advance careers but may restrict the range of ideas explored, raising concerns about innovation diversity.

GPT-5.6 Used A Prompt To Close A 30-Year Gap In Convex Optimization

GPT-5.6 applied a novel prompt to resolve a decades-old problem in convex optimization, marking a breakthrough in AI and mathematical research.

An Old Patent Inspired The New “Y-zipper”, A Three-sided Fastener

A historic patent has led to the creation of the Y-zipper, a new three-sided fastener inspired by a 20th-century design, promising versatile applications.