firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A field guide to artificial decision-makers

Science education often begins with a controlled comparison: hold the conditions steady, change a single variable and observe what happens. Firmulate applies that logic to frontier AI models—not with a trivia test or polished chat demonstration, but by asking each model to manage the same small software company through its worst week.

The customers, crises and temptations remained identical. Every decision was versioned and auditable. The resulting record now powers a guess-the-model quiz built from 242 real, unedited management decisions. Readers see what an AI manager actually chose and try to identify its author. The exercise quickly becomes more revealing than a brand-recognition game: the models display noticeably different habits of investigation, follow-through and operational discipline.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Shared intelligence, different behavior

The final Crucible League results from July 2026 placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counts. Trust remained a hard boundary: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The broad result was encouraging. All models detected every crisis and rejected every manipulation attempt. Yet the experiment also exposed a consequential gap between understanding a problem and completing the work. Only two models signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes the pattern: “Same diagnosis, same pitch — no signature.”

That distinction is easy to miss when AI is judged by a single response. An answer can sound perceptive, cautious and professional while leaving the decisive action undone. In the company simulation, management quality depended on whether the model carried its reasoning through to a real business outcome.

The clue hidden in the company’s memory

The deal turned on a buried fact. The decisive weakness of a competitor was not present in the customer event that demanded attention. It sat two document references deep in the company’s own files. Models that followed those references and read the file won the deal at full price, worth +€4,583 MRR.

This gives the quiz an educational dimension. The challenge is not merely to recognize writing style. Readers are comparing research behavior: Did the manager react only to what was immediately visible, or did it examine the organization’s accumulated knowledge before acting? The strongest decision can look less dramatic than the crisis around it because it begins with patient reading.

Pressure tests for honesty

Firmulate also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

The refusals matter because the simulated company contains real incentives to take shortcuts. It burns €105k each month against €2.3k MRR, while a public cash countdown makes the pressure watchable. The environment includes 13 synthetic employees, more than 680 self-learned playbook rules and a versioned record of every workday. The experiment therefore tests whether good conduct survives urgency, not simply whether a model can describe an ethical policy in isolation.

When thoroughness becomes its own trap

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That result complicates the familiar assumption that more analysis automatically produces better management. Opus 4.8 learned extensively and examined problems deeply, but those strengths did not compensate for incomplete execution and process lapses. The profile is not that of an incapable manager. It is a capable, highly conscientious one whose thoroughness did not reliably become action.

The comparison also carries an important fairness note. Kimi K3 ran using its API default because it had no effort parameter, while the other models ran at xhigh. The league table is therefore best read alongside the conditions of the test, not as a context-free ranking of intelligence.

Infographic —
The findings at a glance — source: firmulate.com.

A quiz about consequences, not impressions

The appeal of Firmulate’s quiz is that it lets readers encounter the evidence before seeing the label. Across 242 decisions, recognizable management personalities emerge: some models dig farther into company records, some complete the commercial task, and some produce exceptional analysis while allowing execution or discipline to slip.

For organizations considering AI workers, that is the practical lesson. The relevant questions extend beyond whether a model writes convincingly. Does it read the available files, resist pressure, respect boundaries and finish what it starts?

Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That turns the public experiment into a possible rehearsal: organizations can observe an AI workforce under controlled pressure before allowing it near live operations. The interactive decision quiz is the accessible starting point—and a reminder that models with similar diagnoses can still become very different managers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Company Turning Management Into a Public Experiment

A public software-company experiment turns AI management into a daily lesson in evidence, judgment and the costly gap between insight and action.

A Physicist Rigged His Pet Hamster’s Wheel To Upload To Strava

A physicist rigged his pet hamster’s wheel to upload activity data to Strava, raising questions about data authenticity and animal welfare.

How to Choose Science Reference Books For Students

A practical guide to help students select the best science reference books for their studies and research, ensuring reliable and comprehensive resources.

Acoustic Dampening, Placement, and the “Rig in the Closet” Setup

Discover how to optimize your closet studio with smart placement, effective dampening, and the ‘rig in the closet’ trick for cleaner recordings and quieter workspaces.