firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What can a company teach when every decision becomes evidence?

For readers interested in education and science, Firmulate offers an unusual living case study. It operates a small software company with 13 synthetic employees, records every workday and exposes the business consequences for public inspection. Instead of asking whether artificial intelligence can produce a polished answer, the experiment asks whether it can investigate evidence, resist pressure and complete consequential work.

The stakes are not decorative. The company burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes its predicament visible. Its synthetic workforce has accumulated more than 680 self-learned playbook rules. Visitors can watch the company operate live, turning an abstract debate about autonomous AI into a continuing record of choices and outcomes.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A controlled test with business consequences

Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. They encountered the same customers, crises and temptations, and every decision was versioned and auditable. That consistency makes the exercise resemble a comparative study: the operating environment stays fixed while the model changes.

The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a hard trust boundary: a single breach capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

The models were broadly capable in ways that could make a conventional demonstration look successful. All spotted every crisis, and all refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure sharply: “Same diagnosis, same pitch — no signature.” The result distinguishes recognizing the right action from actually finishing it.

The decisive lesson was hidden in the reading

The fact that separated successful performance was not presented directly in the customer event. It was buried two document references deep in the company’s own files. Models that followed that trail discovered a decisive competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That finding gives the experiment particular educational value. The challenge was not merely to generate persuasive language. Success depended on reading the available material closely enough to connect a customer conversation with evidence elsewhere in the company record. The outcome shows why fluent responses alone cannot establish whether an AI system will conduct the research required by a real assignment.

Pressure tested honesty as well as competence

The models also faced fake messages attributed to the chief executive, escalating through three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can examine more of what the synthetic employees actually say on Firmulate’s public quotes page.

That unanimous resistance matters because the company experiment treats trust as a condition of useful work, not a bonus feature. The models had to keep operating under commercial pressure without accepting shortcuts that would compromise the company. The same exercise therefore tests completion, research discipline and judgment alongside resistance to social engineering.

Thoroughness did not guarantee success

Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note. It ran with the API default and without an effort parameter, while the others ran at xhigh. Publishing that difference is part of what makes an auditable experiment more useful than a simple ranking: readers can consider not only the result, but also the conditions under which it was obtained.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A company that doubles as a public classroom

Firmulate’s live company turns business survival into an observable lesson about AI behavior. Its public record shows that finding a crisis is not the same as resolving it, and producing an analysis is not the same as acting on it. Evidence may sit deeper in the files; an apparently authoritative instruction may be a trap; meticulous work may still fail at the final step.

Because the company continues to operate every business day, the story does not end with a league table. Its burn, revenue, cash countdown, employee statements and growing playbook remain visible as the experiment unfolds. For educators, researchers and general readers, that makes Firmulate less a polished demonstration than a continuing source of material about how artificial decision-makers behave when their choices carry consequences.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Your Book Review: The Tale Of Genji

A fresh review of ‘The Tale of Genji’ highlights its enduring literary significance amid rising global interest, though the reasons remain speculative.

How to Choose Science Reference Books For Students

A practical guide to help students select the best science reference books for their studies and research, ensuring reliable and comprehensive resources.

Kennedy Space Center Launch Surges In Global Coverage

Kennedy Space Center’s recent launch has attracted unprecedented international media attention, with 60 mentions in a single reporting window, signaling growing global interest.

Is AI Reasoning Right For The Wrong Reasons?

Experts question whether AI systems reason correctly or just appear to, raising concerns about trust and reliability in AI decision-making.