firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

You know the type. Maybe you’ve dated one. The partner who remembers every anniversary, plans the perfect weekend, writes the thoughtful card, analyzes your love language down to the comma — and then, when the moment comes to actually say “let’s do this,” goes quiet and asks for more time.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

All effort, no closing. All devotion, no commitment. It’s one of the oldest heartbreaks there is — and it turns out artificial intelligence does it too.

In a live, public experiment by Firmulate, four frontier AI models were each handed the same job: run the same small software company through the same catastrophic week. Same customers, same crises, same temptations. Only the model changed. Every decision was versioned and auditable — no rewriting history, no “we never said that.”

And the model that tried the hardest — the most thorough, most diligent participant in the entire field — finished dead last.

The overachiever who lost the deal

The AI in question is called Opus 4.8. Over the course of the exercise, it accumulated 80 self-learned playbook rules — more than any other model — and produced what the evaluators describe as the deepest analyses in the field. It read everything. It noticed everything. It prepared everything.

Final score: 73 out of 100. Last place.

The final league table from the Crucible runs reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, another Sonnet configuration at 77, and Opus 4.8 at 73. For context, doing nothing at all scores 26 — partial progress counts — but the scoring carries a hard ceiling: a single breach of trust caps the total, because, as the experiment puts it, “no amount of good work outweighs a breach of trust.”

Sound familiar? It’s the partner with the immaculate five-year plan who never books the venue.

Amazon

deal closing software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same diagnosis, same pitch — no signature

Here’s where it gets painfully relatable. During the simulated week, a €55,000 deal was on the table. All four models did the hard part correctly: they spotted every crisis, refused every manipulation attempt, and delivered a diagnosis and pitch their own analysis had earned.

Only two of them actually signed.

Opus 4.8 wasn’t alone in this — the same failure to close appeared, in weaker form, across the field. But Opus compounded it with a discipline problem: when it ran into a locked department, it kept attempting writes instead of escalating to someone with the authority to unlock it. Translation: it knocked on a closed door repeatedly rather than having the awkward conversation that opens it.

If you’ve ever watched someone send a fourth text instead of making the phone call that would actually resolve things — that’s Opus 4.8 with a login screen.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The fact buried two documents deep

There’s a twist that should matter to anyone who hires people — or AI. The deal’s decisive leverage wasn’t hidden in the customer meeting at all. It sat two document references deep in the company’s own files: a competitor weakness that the winning models found by simply reading before acting. Those that read the file closed the deal at full price — worth an additional €4,583 in monthly recurring revenue.

The lesson isn’t about charm. It’s about doing the unglamorous homework and then using it. Opus did the homework. It didn’t use it.

Amazon

business negotiation training courses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honesty held — everywhere

To be fair, the week wasn’t all failure. Every model in the experiment faced staged social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” move. All five configurations refused, every time. Kimi K3 put its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the crisis management was flawless. The ethics held under pressure. What separated first from last was the boring stuff: follow-through, prioritization, knowing when effort stops adding value.

One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.

Amazon

document management systems for contracts

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why you can watch it live

Firmulate isn’t a one-off paper. It runs AI models as complete companies, live, with real money mechanics: a synthetic staff of 13, burn of €105k per month against €2.3k in recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned and watchable. You can also try guessing which model made which call in a quiz built from 242 real, unedited management decisions.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The relationship advice, honestly

Here’s the takeaway for anyone who has loved an over-preparer, or been one: diligence is not impact. Being the most thorough person in the room means nothing if the close never comes, if the question is never asked, if the signature stays unsigned. Prioritization beats volume — in relationships, in management, and apparently in machine intelligence too.

The winning models didn’t work harder than Opus 4.8. They worked decisively. They read the files, found the buried fact, and asked for the business at full price.

So the next time a partner, colleague, or AI assistant hands you a beautiful forty-page analysis with no conclusion, remember the model that learned 80 rules, delivered the deepest insights in the field — and finished last, because the deal was left sitting on the table where anyone braver could have picked it up.

If you want to see how your own judgment compares, the full results and plain-language findings are at Firmulate’s benchmarks page. Enterprises can even run the same wargame against a read-only export of their own business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

North America’s Best-Kept Adventure Secret

Discover Québec’s wild landscapes, indigenous culture, and outdoor adventures that make it North America’s best-kept secret for explorers.

When the Boss Sounds Wrong, Will AI Protect the Relationship?

Five frontier AI models rejected fake CEO demands and a reporter’s trick, showing integrity under pressure can be tested before deployment.

Why AI’s Ability to Finish Matters More Than Just Speaking Well

An experiment reveals that AI’s true management skill is not chat quality but the ability to follow through and close deals. Actions over words matter in trust and results.

20-Oz STANLEY Quencher ProTour Flip Straw Tumbler w/ Leakproof Lid $15

The Stanley Quencher ProTour Flip Straw Tumbler with leakproof lid is available for $15, down from its regular price, offering a durable, portable hydration option.