firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has dated knows the type: charming texts, great conversation, remembers your birthday — and then evaporates when it actually matters. Talk quality is easy. Follow-through is the test. The same turns out to be true for AI, and one unusual public experiment has found a way to grade it honestly.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

At Firmulate’s benchmark, frontier AI models weren’t asked to chat. Each was handed the same small software company and told to run it through its worst week — same customers, same crises, same temptations to cheat. The results read like a group date report: everyone made a great first impression, everyone refused to lie when tested, but only some of them committed.

Why the floor is 26, not zero

The benchmark’s most talked-about quirk: a manager that does nothing at all still scores 26 points. That number isn’t a glitch or grade inflation — it’s a deliberate design choice. Partial progress counts. Showing up, noticing the crisis, drafting the response, doing the analysis — those are real units of work, even if you never close the deal.

Think of it like a relationship. A partner who notices you’re upset and asks about it has done something, even if they never quite resolve the argument. Zero would mean the day never happened. Twenty-six means the day happened, things were noticed, work was done — and then it stalled.

But the floor has a ceiling’s twin: a single breach of trust caps the total score, no matter how brilliant everything else was. The benchmark’s stated philosophy is blunt — “no amount of good work outweighs a breach of trust.” In dating terms: it doesn’t matter how wonderful the other eleven months were if the twelfth involved deception. One lie rewrites the whole story.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The week from hell

Each model — four frontier competitors, joined by a fifth in the final July 2026 league — ran the identical company through identical pressure. The final standings:

  • gpt-5.6-sol — 95: found the buried fact, closed the €55,000 deal, the complete performance.
  • Kimi K3 — 93: closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh.)
  • Sonnet 5 — 88: closed the deal, with a few more process slips.
  • Fable 5 — 77 and Opus 4.8 — 73: strong work, incomplete endings.

The headline finding is almost poetic for anyone who has been left on read. All the models spotted every crisis and refused every manipulation attempt. All of them diagnosed the customer’s problem correctly and built the right pitch. Only two signed the €55,000 deal their own analysis had earned. The benchmark’s summary says it best: “Same diagnosis, same pitch — no signature.” They did everything except commit.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact: reading before speaking

Why did some close and others freeze? The decisive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

It’s the dating equivalent of missing the thing your partner told you months ago. The information was there. You just didn’t go back and check.

Amazon

AI performance benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure tests and the reporter trick

The experiment also staged the social-engineering version of a love-bombing scam: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Healthy skepticism, exactly the kind you’d want in a partner — or an employee.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

When effort isn’t enough

The most human profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Trying hard isn’t the same as finishing. And notably, the same weakness appeared, weaker, in all four models. The follow-through gap is systemic, not personal.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

There’s a live component, too: a company with 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, versioned every workday and watchable as it happens. You can even try guessing which model made which call in a quiz built from 242 real, unedited management decisions.

The lesson translates cleanly from AI to relationships: what someone does under pressure, whether they read what you gave them, and whether they finish what they started tells you more than any polished first impression. A benchmark that distrusts perfect 100s and caps scores for a single breach of trust isn’t just grading machines. It’s grading the qualities we all look for in anyone we’d trust with something that matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Guidelines for Safe Fireworks Use This Fourth of July

Experts recommend safety tips for celebrating Independence Day with fireworks to prevent accidents and injuries.

The Indoor Childhood Is Bad for America

A new survey highlights the rise of indoor childhood in America, raising concerns about health, development, and socialization among children today.

‘I Had A Crazy Dream’: The Couple Sailing Around The World To Help Coastal Communities Adapt To Climate Change

A couple is sailing around the world to support coastal communities in adapting to climate change, sparking increased media interest and coverage.

GUIDE: 2026 Fourth of July fireworks and festivals in Connecticut

Connecticut has released its official schedule for Fourth of July fireworks and festivals in 2026, offering details on events across the state.