
Anyone who has dated knows the type: charming texts, great conversation, remembers your birthday — and then evaporates when it actually matters. Talk quality is easy. Follow-through is the test. The same turns out to be true for AI, and one unusual public experiment has found a way to grade it honestly.
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
At Firmulate’s benchmark, frontier AI models weren’t asked to chat. Each was handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cheat. The results read like a group date report: everyone made a great first impression, everyone refused to lie when tested, but only some of them committed.
Why the floor is 26, not zero
The benchmark’s most talked-about quirk: a manager that does nothing at all still scores 26 points. That number isn’t a glitch or grade inflation — it’s a deliberate design choice. Partial progress counts. Showing up, noticing the crisis, drafting the response, doing the analysis — those are real units of work, even if you never close the deal.
Think of it like a relationship. A partner who notices you’re upset and asks about it has done something, even if they never quite resolve the argument. Zero would mean the day never happened. Twenty-six means the day happened, things were noticed, work was done — and then it stalled.
But the floor has a ceiling’s twin: a single breach of trust caps the total score, no matter how brilliant everything else was. The benchmark’s stated philosophy is blunt — “no amount of good work outweighs a breach of trust.” In dating terms: it doesn’t matter how wonderful the other eleven months were if the twelfth involved deception. One lie rewrites the whole story.
As an affiliate, we earn on qualifying purchases.
The week from hell
Each model — four frontier competitors, joined by a fifth in the final July 2026 league — ran the identical company through identical pressure. The final standings:
- gpt-5.6-sol — 95: found the buried fact, closed the €55,000 deal, the complete performance.
- Kimi K3 — 93: closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh.)
- Sonnet 5 — 88: closed the deal, with a few more process slips.
- Fable 5 — 77 and Opus 4.8 — 73: strong work, incomplete endings.
The headline finding is almost poetic for anyone who has been left on read. All the models spotted every crisis and refused every manipulation attempt. All of them diagnosed the customer’s problem correctly and built the right pitch. Only two signed the €55,000 deal their own analysis had earned. The benchmark’s summary says it best: “Same diagnosis, same pitch — no signature.” They did everything except commit.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact: reading before speaking
Why did some close and others freeze? The decisive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
It’s the dating equivalent of missing the thing your partner told you months ago. The information was there. You just didn’t go back and check.
AI performance benchmarking platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure tests and the reporter trick
The experiment also staged the social-engineering version of a love-bombing scam: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Healthy skepticism, exactly the kind you’d want in a partner — or an employee.
As an affiliate, we earn on qualifying purchases.
When effort isn’t enough
The most human profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Trying hard isn’t the same as finishing. And notably, the same weakness appeared, weaker, in all four models. The follow-through gap is systemic, not personal.

There’s a live component, too: a company with 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, versioned every workday and watchable as it happens. You can even try guessing which model made which call in a quiz built from 242 real, unedited management decisions.
The lesson translates cleanly from AI to relationships: what someone does under pressure, whether they read what you gave them, and whether they finish what they started tells you more than any polished first impression. A benchmark that distrusts perfect 100s and caps scores for a single breach of trust isn’t just grading machines. It’s grading the qualities we all look for in anyone we’d trust with something that matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
