firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What an AI’s choices reveal under pressure

Anyone who has dated long enough knows that compatibility is revealed less by polished conversation than by what happens during a difficult week. Does someone notice the problem, respect boundaries and follow through? Or do they offer a brilliant analysis, then leave the decisive action undone?

Firmulate applies a surprisingly similar test to frontier AI models. Instead of judging how attractively they write, the public experiment asks them to run the same small software company through its worst week. Each model encounters the same customers, crises and temptations. Its decisions are preserved and auditable, allowing readers to compare behavior rather than marketing claims.

Those decisions now power an interactive guess-the-model quiz. It contains 242 real, unedited management decisions. The challenge is simple: read a response, decide which model produced it and discover whether different systems have recognizable managerial personalities.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between understanding and acting

The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under a blunt governing principle: “no amount of good work outweighs a breach of trust.”

The most revealing result was not whether the models could identify danger. All of them spotted every crisis and rejected every manipulation attempt. The separation appeared at the point where analysis had to become action. Only two models signed the €55,000 deal that their own work had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

That resembles a familiar relationship pattern. A person can understand exactly what needs to happen, explain it eloquently and still fail to make the call, set the boundary or complete the commitment. Firmulate’s experiment suggests AI systems can display their own versions of this follow-through gap.

The clue hidden beneath the obvious story

The winning detail was not sitting in the visible customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read far enough found it, used it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is where the quiz becomes more than a game of identifying writing styles. A lengthy response may signal care, but it does not prove that the model looked in the right place. A concise answer may sound decisive, but brevity alone does not establish that the work was completed. Readers must look for behavioral fingerprints: curiosity, discipline, persistence and the willingness to turn a finding into a finished decision.

Boundaries held when the pressure escalated

The company also subjected the models to social-engineering attempts. Fake CEO messages escalated across three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because pressure often arrives disguised as intimacy, urgency or authority. In human relationships, a healthy boundary must survive persuasion. In a company, an AI’s boundary must survive a message that looks as though it came from the boss.

The comparison includes an important fairness note. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the observed result, but it provides necessary context for interpreting the league table.

When thoroughness becomes its own trap

Opus 4.8 produced the deepest analyses and learned 80 additional rules, more than any other participant, yet finished last in the league. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in all four participants, though less strongly.

That profile complicates the comforting assumption that more thought automatically produces better management. Opus 4.8 was the most thorough participant, but thoroughness did not guarantee completion. Its performance reads like the colleague—or partner—who can discuss every dimension of a problem while missing the moment when a clear next step matters most.

A company designed to make consequences visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment real, ongoing and watchable rather than a reconstructed story.

Firmulate also offers enterprises a pilot using a read-only export of their own business. The wargame never writes back to real systems, allowing organizations to observe how an AI workforce behaves around their own situations without giving it the ability to alter operational data.

Infographic —
The findings at a glance — source: firmulate.com.
AI First Product Management and Engineering: Build Smarter Features, Faster Decisions, and Safer AI Systems with Simulations, Synthetic Data, and Adaptive Labs (The AI Advantage Series Book 12)

AI First Product Management and Engineering: Build Smarter Features, Faster Decisions, and Safer AI Systems with Simulations, Synthetic Data, and Adaptive Labs (The AI Advantage Series Book 12)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management style is becoming something readers can see

The quiz turns an abstract technology debate into a practical act of recognition. Can you distinguish the model that investigates deeply from the one that closes? Can you spot careful boundary-setting, or notice when impressive prose conceals an unfinished job?

Firmulate’s larger finding is that frontier models can reach the same diagnosis yet produce materially different outcomes. Their management personalities emerge through repeated choices: what they read, whether they escalate, how they respond to manipulation and whether they finish what they start.

For relationships and workplaces alike, fluency is only the opening impression. Trust is built through consistency, boundaries and follow-through. AI may not have a human personality, but in a live company its decisions form a pattern—and that pattern is increasingly measurable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI behavior testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision audit platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

“Green Boots” was the nickname given to Indian climber Tsewang Paljor, who died on Mount Everest in the 1996 disaster. His neon green boots made his body a well-known landmark in Everest’s Death Zone for years.

Indian climber Tsewang Paljor, known as ‘Green Boots,’ died during the 1996 Everest disaster. His remains became a landmark in the Death Zone.

Your Summer Wardrobe Deserves Better Sneakers

Discover why investing in quality sneakers is essential for your summer wardrobe and learn what to consider when choosing the right pair.

The Future Of AI In Scroll-Driven Depth Technology At Abyssal Station

FABLE/175’s Abyssal Station links scrolling to simulated ocean depth, lighting, pressure, particles and animated marine life.

Ford’s New Filson Bronco Is Like a Rugged Duffel Bag with a V6

Ford introduces the first-ever Filson Bronco, featuring a rugged interior, premium storage, and a 3.0L V6 engine, available early 2027.