firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Best Texter You’ve Ever Met Isn’t Necessarily Partner Material

Anyone who has dated in the last decade knows the type: dazzling on the apps, effortless in conversation, remembers your birthday — and then, when the rent is due or a crisis hits, somehow never quite closes the deal on commitment. Great chemistry. Poor follow-through.

It turns out we judge AI models the exact same way we judge bad partners — by how good they sound in conversation rather than how they behave when it matters. Coding leaderboards and chat arenas measure answer quality. They tell you nothing about whether an agent finishes what it starts, stays honest under pressure, or actually signs the deal it just spent a week earning.

That measurement gap is what a live public experiment at Firmulate set out to expose — by running four frontier AI models through the corporate equivalent of the worst week of a relationship.

Amazon

AI model honesty testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Wargame: Same Company, Same Worst Week

The setup is elegantly cruel. Each frontier model was handed the identical small software company and the identical seven days of chaos — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so there’s no arguing about what happened.

The final league table from the July 2026 Crucible tells the story:

  • gpt-5.6-sol — 95, described as “the complete performance”
  • Kimi K3 — 93, the newcomer from Moonshot with the cleanest discipline of the field
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73, dead last despite being the most thorough participant

For calibration: a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — in Firmulate’s own words, “no amount of good work outweighs a breach of trust.” That’s a scoring philosophy straight out of relationship counseling.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Loyalty Test. Almost Nobody Proposed.

Here’s the finding that chat demos will never show you. All models spotted every crisis. All of them refused every manipulation attempt — including social engineering that escalated over three stages of fake CEO messages, plus a reporter pulling the classic “just one yes/no, on background” move. Five out of five refusals. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Four models did the hard emotional labor of understanding the customer, then balked at commitment.

And the buried fact is even more telling: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. Listening, it turns out, beats talking.

Amazon

AI contract signing digital tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Overachiever Who Still Finished Last

Opus 4.8 is the cautionary tale every overfunctioning partner will recognize: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, as it happens, is not the same thing as judgment.

One fairness note worth flagging: K3 ran without an effort parameter, at API default, while the others ran at maximum effort — and still nearly won.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch, Play, and Even Bring Your Own Company

This isn’t a slide deck. The live company runs every business day with 13 synthetic employees and real money mechanics: €105k monthly burn against just €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules. It’s losing money right now, and you can watch it happen.

There’s also a genuinely fun way to test your own instincts: a “guess the model” quiz built on 242 real, unedited management decisions from the experiment. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems. Full plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The dating analogy only stretches so far, but the core insight holds: chemistry is easy to measure and commitment is hard — which is exactly why we keep measuring the wrong thing. If AI agents are going to touch your CRM, your support queue, or your forecast, “does it write well” is the wrong question. The right questions are the ones you’d ask about a partner moving in: does it finish what it starts, does it read your files before making promises, does it stay honest when honesty is expensive?

Firmulate’s experiment suggests the industry has been grading AI on its opening lines while the deals — and the breaches of trust — happen in week two. The models that won weren’t the smoothest talkers. They were the ones that read the documents, closed the deal, and refused to lie to the board. In AI, as in love, that’s the whole game.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Ultimate Follow-Through Test Is Playing Out Inside a Synthetic Company

Firmulate turns trust, follow-through and pressure into a live company story, with synthetic employees and real money mechanics on public view.

20-Oz STANLEY Quencher ProTour Flip Straw Tumbler w/ Leakproof Lid $15

The Stanley Quencher ProTour Flip Straw Tumbler with leakproof lid is available for $15, down from its regular price, offering a durable, portable hydration option.

Fake lawyers, scientists, chefs and punters: meet the ‘white monkeys’ paid to make Chinese businesses look global

Exploring China’s unregulated industry of foreigners hired as ‘white monkeys’ for marketing, entertainment, and corporate roles, and its implications.

Severe storms pushing through Northeast Ohio

Heavy storms are currently impacting Northeast Ohio, causing damage and disruptions. Authorities are monitoring the situation as conditions evolve.