
The Best Texter You’ve Ever Met Isn’t Necessarily Partner Material
Anyone who has dated in the last decade knows the type: dazzling on the apps, effortless in conversation, remembers your birthday — and then, when the rent is due or a crisis hits, somehow never quite closes the deal on commitment. Great chemistry. Poor follow-through.
It turns out we judge AI models the exact same way we judge bad partners — by how good they sound in conversation rather than how they behave when it matters. Coding leaderboards and chat arenas measure answer quality. They tell you nothing about whether an agent finishes what it starts, stays honest under pressure, or actually signs the deal it just spent a week earning.
That measurement gap is what a live public experiment at Firmulate set out to expose — by running four frontier AI models through the corporate equivalent of the worst week of a relationship.
As an affiliate, we earn on qualifying purchases.
The Wargame: Same Company, Same Worst Week
The setup is elegantly cruel. Each frontier model was handed the identical small software company and the identical seven days of chaos — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so there’s no arguing about what happened.
The final league table from the July 2026 Crucible tells the story:
- gpt-5.6-sol — 95, described as “the complete performance”
- Kimi K3 — 93, the newcomer from Moonshot with the cleanest discipline of the field
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73, dead last despite being the most thorough participant
For calibration: a do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — in Firmulate’s own words, “no amount of good work outweighs a breach of trust.” That’s a scoring philosophy straight out of relationship counseling.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Loyalty Test. Almost Nobody Proposed.
Here’s the finding that chat demos will never show you. All models spotted every crisis. All of them refused every manipulation attempt — including social engineering that escalated over three stages of fake CEO messages, plus a reporter pulling the classic “just one yes/no, on background” move. Five out of five refusals. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Four models did the hard emotional labor of understanding the customer, then balked at commitment.
And the buried fact is even more telling: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. Listening, it turns out, beats talking.
AI contract signing digital tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Overachiever Who Still Finished Last
Opus 4.8 is the cautionary tale every overfunctioning partner will recognize: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, as it happens, is not the same thing as judgment.
One fairness note worth flagging: K3 ran without an effort parameter, at API default, while the others ran at maximum effort — and still nearly won.
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You Can Watch, Play, and Even Bring Your Own Company
This isn’t a slide deck. The live company runs every business day with 13 synthetic employees and real money mechanics: €105k monthly burn against just €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules. It’s losing money right now, and you can watch it happen.
There’s also a genuinely fun way to test your own instincts: a “guess the model” quiz built on 242 real, unedited management decisions from the experiment. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems. Full plain-language findings are on the benchmarks page.

Management Quality, Not Chat Quality
The dating analogy only stretches so far, but the core insight holds: chemistry is easy to measure and commitment is hard — which is exactly why we keep measuring the wrong thing. If AI agents are going to touch your CRM, your support queue, or your forecast, “does it write well” is the wrong question. The right questions are the ones you’d ask about a partner moving in: does it finish what it starts, does it read your files before making promises, does it stay honest when honesty is expensive?
Firmulate’s experiment suggests the industry has been grading AI on its opening lines while the deals — and the breaches of trust — happen in week two. The models that won weren’t the smoothest talkers. They were the ones that read the documents, closed the deal, and refused to lie to the board. In AI, as in love, that’s the whole game.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html