firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

When the Answer Isn’t in the Question

Anyone who has written a research paper knows the drill: the claim sits in one source, but the evidence that makes it true sits in that source’s source. Teachers call it synthesis. Librarians call it following the citation chain. In AI research, it goes by a drier name — the multi-hop question — and it has long been the gap between models that recite and models that reason.

A new public experiment from Firmulate, which runs frontier AI models as complete companies through identical worst-case weeks, has just turned that academic concern into a hard commercial number. Buried two document references deep in a fictional software firm’s own files was a single decisive fact — a competitor’s weakness. The AI models that followed the citation chain found it, used it, and closed a €55,000 deal at full price. The ones that stopped one hop short didn’t just answer less precisely. They lost the sale automatically.

Amazon

AI reading comprehension software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Only the Model Changes

The setup is elegant in the way good experimental design always is. Each frontier model was handed the same small software company and the same brutal week: identical customers, identical crises, identical temptations to cut corners. Every decision was versioned and auditable, so the runs can be compared line by line. The final league table from July 2026 puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline still scores 26 — partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.

The headline finding cut against expectations. Every model spotted every crisis. Every model refused every manipulation attempt. But only two of the five signed the €55,000 contract that their own analysis had earned. The experiment’s summary puts it bluntly: same diagnosis, same pitch — no signature.

The Fact That Was Two Documents Deep

Why did three models come so close and stop? The answer traces back to a single buried fact. The decisive competitor weakness was not in the customer’s event log, where the models were all looking. It sat two document references deep in the company’s own files — a fact you could only reach by reading one document, noticing its reference to another, and going to read that too.

That is a multi-hop retrieval problem wearing a business suit. And its value was quantified: models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that didn’t, lost it. Reading comprehension, it turns out, is a purchase-deciding property of an AI agent — as measurable as latency or cost per query.

Pressure, Trickery, and One Model’s On-Record Reasoning

The week also included social engineering: fake CEO messages that escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 for 5. Kimi K3’s refusal came with reasoning worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then there is the Opus 4.8 paradox, which will feel familiar to anyone who has known a brilliant student who over-studies and under-finishes. It was the most thorough participant in the field — 80-plus self-learned playbook rules added, the deepest analyses of any model — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four of the other models.

An Open Laboratory, Not a Lab Report

Firmulate’s live company is watchable in real time: 13 synthetic employees, real money mechanics with a burn of €105k per month against €2.3k in MRR, a public cash countdown, and 680-plus self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day as new benchmark runs finish.

One fairness caveat the experiment publishes itself: Kimi K3 ran at its API-default effort setting while the other models ran at maximum effort — and still placed second. And for readers who want to test their own judgment, 242 real, unedited management decisions from the runs power a guess-the-model quiz.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Lesson for Anyone Teaching or Testing AI

If you evaluate AI systems — or teach students who will — the Firmulate results argue for a specific shift in what we measure. Fluency is now cheap; every model wrote a convincing pitch. Honesty under pressure held up across the board. What separated first place from last was the unglamorous, scholarly habit of following references to their source, and the follow-through to finish what an analysis started.

In other words, the old library skill — check the citation’s citation — is now worth exactly €55,000 in a single deal, and it shows up in a public league table anyone can audit. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Before you hand an AI agent your CRM, support queue, or forecast, the question is no longer whether it writes well. It’s whether it does its homework.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI-generated videos to maximally drive a target brain region

Researchers develop AI-generated videos designed to target and stimulate specific brain regions, advancing neuroscience and potential therapies.

Improving Heuristics For A* Pathfinding

New heuristics enhance A* algorithm efficiency, promising faster pathfinding in robotics and gaming applications.

Japan’s Hayabusa2 Probe To Conduct Flyby Of Torifune Asteroid

Japan’s Hayabusa2 spacecraft is set to fly by the Torifune asteroid for scientific observations, marking a new phase in its mission to study near-Earth objects.

Best Science Reference Books For Students Compared

Compare popular science reference books for students to find the best fit for learning, budget, and study needs. Make an informed choice today.