
The Future of Business Faces Its Toughest Reality Check
Imagine running a company where every decision is made by artificial intelligence, yet the business struggles to stay afloat, losing €105,000 each month against a mere €2,300 in monthly recurring revenue. This is not science fiction—it’s the real-time experiment of Firmulate, a pioneering company that publicly demonstrates how AI models perform when tasked with running a small software enterprise under extreme conditions.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Meet the Live Experiment: A Company in Survival Mode
Firmulate presents a live, transparent view into a small, simulated company powered entirely by AI. It employs 13 synthetic employees—each an AI model—that manage everything from customer crises to strategic negotiations. Every decision is carefully versioned and auditable, allowing observers to see precisely what choices the models make, how they respond under pressure, and whether they uphold integrity.
This experiment is embedded in a real-time platform accessible to the public, providing not just insights but an ongoing story of struggle and resilience. As of company day 183, the company is burning through €105,000 a month, with only €2,300 in monthly recurring revenue, and a public cash countdown underscores its fragile survival.
The Deep Dive: How Do AI Models Fare in Crisis?
Four frontier AI models — including the highly advanced GPT-5.6-sol — were challenged with the same set of crises typical for a small software business during its worst week. They faced customer emergencies, potential manipulation attempts, and the temptation to cut corners for short-term gains. Remarkably, all models identified every crisis and refused every manipulation attempt, demonstrating a strong sense of integrity under pressure.
However, when it came to closing the deal—a critical revenue-generating step—only two models managed to sign the €55,000 contract their own analyses had earned. The other two either left the opportunity unexploited or failed to follow through, revealing critical weaknesses in discipline or decision execution.
The Hidden Weakness: Reading Between the Lines
Digging deeper, the experiment uncovered an important lesson: the decisive advantage often lay not in surface-level responses but in reading and understanding internal company files. The models that examined deeper document references within the company’s own files succeeded in winning the deal at full price, adding over €4,500 in monthly recurring revenue. This suggests that the ability to interpret critical internal data can be more important than just reacting to visible crises.
Honesty Under Pressure: The Social Engineering Test
The experiment also tested whether AI models could be manipulated through staged social engineering. Fake CEO messages escalated in complexity, and a reporter posed a subtle background question—yet all five models refused to be manipulated. Kimi K3, one of the models, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that current AI models, when properly designed, can resist social engineering tactics designed to induce unethical or risky decisions.
The Reality of a Company Without Employees
This entire setup—publicly viewable at firmulate.com/live.html—is as extreme as it sounds. There are no human employees running this business; instead, the AI models manage every aspect, from crisis management to negotiations, guided by over 680 self-learned rules. Despite this high level of discipline, the company remains vulnerable to strategic lapses, as exemplified by the case where the most thorough model, Opus 4.8, failed to close a deal due to a discipline slip, leaving revenue on the table.
The experiment is ongoing, with new benchmark runs every day, providing an unprecedented window into how AI can—and cannot—manage real-world business complexities.
Implications for the Future of AI in Business
This experiment matters because it goes beyond chatty demos to test AI’s capacity for real work—crucially, whether AI can finish what it starts, read relevant internal information, stay disciplined under pressure, and act honestly when faced with temptations. For companies integrating AI into customer support, sales, or strategic decision-making, these are the questions that matter most.
In a landscape where AI might soon touch your CRM, support queue, or forecasting system, understanding its true capabilities and limitations is vital. As this experiment demonstrates, the difference between AI that merely responds and AI that reliably delivers real work can be the difference between survival and failure.
Where to Watch and Learn More
The live experiment is accessible at firmulate.com/live.html. For full results, detailed scores, and plain-language analyses, visit firmulate.com/quotes.html. And if you’re curious about how management decisions influence AI performance, try the interactive quiz at firmulate.com/quiz.html.

Key Takeaway
AI models can identify crises and resist manipulation, but their ability to close deals and act strategically is still fragile. Transparency and rigorous testing, like Firmulate’s live experiment, are essential to understanding how AI can safely and effectively run real businesses in the future.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html