firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Science classes teach the difference between knowing and doing — reciting the textbook answer versus running the experiment. A public, still-running experiment at Firmulate has now demonstrated that gap inside AI systems, with unusual rigor: four frontier AI models were each handed the same small software company and asked to steer it through its worst week. All of them diagnosed every crisis correctly. All of them refused every attempt at manipulation. But only two of them actually signed the €55,000 deal their own analysis had earned.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

For educators and researchers, the setup will feel familiar: a controlled experiment, a single variable, and a measurable outcome. The variable was the AI model. Everything else — customers, crises, temptations, the seed — was identical. The result is a benchmark that measures management quality, not chat quality.

The Experiment

Each frontier model ran the same small software company through the same brutal week: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, like lab notes you can replay.

The final league standings from July 2026:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For scale, the do-nothing baseline scores 26. Partial progress counts toward a score, but a single breach of trust caps the total — the experiment’s designers phrase it as “no amount of good work outweighs a breach of trust.” Notably, Kimi K3 achieved second place while running at its API-default effort setting, while the other models ran at xhigh effort.

The Finding: Knowledge Without the Signature

The headline result is a behavioral gap invisible in chat demos. All models spotted every crisis. All refused every manipulation attempt. Yet only two closed the deal they had correctly diagnosed and correctly pitched — “Same diagnosis, same pitch — no signature.”

The buried fact is even more instructive: the decisive competitive weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson maps neatly onto education research: deep reading of source material, not skimming the headline, is what produced the outcome.

Social Engineering: Five Out of Five Said No

The experimenters also staged impersonation attacks — fake CEO messages escalating over three stages, plus a reporter’s trick (“just one yes/no, on background”). All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That consistency is a genuinely encouraging data point for anyone worried about AI agents in sensitive workflows.

The Thoroughness Trap

Opus 4.8 is the case study: it was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses — and still finished last. It left the close on the table, and its discipline slipped: it attempted writes into a locked department rather than escalating. The same weakness appeared, more weakly, in all four models. Diligence, it turns out, is not the same as judgment.

Still Running, Still Watchable

The live company is real and observable at firmulate.com: 13 synthetic employees, real money mechanics — burn of €105k/month against €2.3k MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. A companion quiz at firmulate.com/quiz.html lets anyone try to guess which model made which of 242 real, unedited management decisions.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The scientific takeaway for a general audience: we now have a repeatable, auditable way to study AI decision-making under pressure — not by asking models questions, but by making them run a company with real consequences attached. And the results are humbling for the models and instructive for us: crisis detection and integrity are largely solved at this level; initiative, closure, and reading the fine print two references deep are where the winners separated from the losers.

For enterprises, the experiment moves from observation to application: you can run the same wargame against a read-only export of your own business — your customers, your pipeline, your rules — with churn waves, price increases, competitor attacks, PR crises, and social-engineering pressure. Nothing ever writes back to real systems, and you receive a board report with the model ranking and the weak points of your own playbooks.

Ready to stress-test your own company before reality does it for real? Explore the enterprise pilot at firmulate.com/pilot.html or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Educational Science Kits For Kids: A Back to school Guide

Discover how educational science kits for kids spark curiosity, teach core concepts, and keep children engaged. Find the best kits for all ages today!

How to Choose Science Reference Books For Students

A practical guide to help students select the best science reference books for their studies and research, ensuring reliable and comprehensive resources.

Fermat’s Last Theorem In Lean 4

Mathematicians have completed a formal proof of Fermat’s Last Theorem using Lean 4, marking a significant milestone in formal verification and mathematical proof systems.

Stop Anthropomorphizing Intermediate Tokens As Reasoning/Thinking Traces (2025)

Researchers warn against interpreting intermediate AI tokens as evidence of reasoning, emphasizing the need for clearer evaluation methods in AI development.