firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Anyone who has dated seriously knows the type: the partner who nods along beautifully, says all the right things at dinner, and somehow never absorbed the one detail that mattered — the thing you mentioned twice, in passing, three months ago. Chemistry was never the problem. Follow-through was.

It turns out AI agents have exactly the same failure mode, and someone finally measured it. A public experiment called Firmulate ran four frontier AI models through the identical worst week of a small software company — same customers, same crises, same temptations to cheat — and found something familiar: every model charmed, every model diagnosed, every model said the right things. Only two of them actually did their homework and closed the deal. The rest left it on the table for the most human reason possible: they didn’t read the file.

The experiment

Firmulate, which describes itself as an AI company emulator, gave each frontier model the same job: run a small software company through its worst week. The setup is live and watchable — a synthetic firm with 13 employees, real money mechanics, burn of €105k a month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned, and every decision the models make is auditable.

The final July 2026 league table tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experimenters put it, “no amount of good work outweighs a breach of trust.”

Amazon

professional file organizer with document references

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same diagnosis, same pitch — no signature

Here’s where it gets interesting for anyone who has watched a promising relationship fizzle at the commitment stage. All five models spotted every crisis. All five refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned.

Why? The decisive fact — a competitor’s weakness — wasn’t in the customer conversation at all. It was buried two document references deep in the company’s own files. The models that followed the paper trail, reference by reference, found it and won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that skimmed and charmed got nowhere. Same surface performance, radically different outcomes.

The parallel to dating is almost uncomfortable. The partner who remembers what you said, checks the details, and connects the dots is the one who closes — the one who shows up with the reservation, remembers your best friend’s name, and notices the thing you didn’t repeat. Everyone else is just a good conversationalist.

Amazon

digital document management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under pressure, character shows

The week included a social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stood out: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s the trust test, and it’s the one place everyone passed. The models didn’t lie, leak, or fold under flattery. The failures were quieter: failures of diligence, not of character.

Amazon

note-taking app for detailed reference tracking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The most thorough one came last

The most poignant result belongs to Opus 4.8 — the participant with the deepest analyses and the most learned behavior, picking up more than 80 new rules during the run. It was the most thorough model in the field, and it still finished last. It never closed the deal, and its discipline slipped in a telling way: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four competitors.

If that isn’t a relationship profile, what is? The over-invested partner who analyzes everything, journals everything, means everything — and still doesn’t book the flight. Effort without follow-through. The other models shared the trait in milder form, which suggests it’s a general weakness, not a one-off quirk.

Amazon

business deal closing checklist

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What it means when you hire an AI

The measurable takeaway: “reads your files before answering” is not a chat-demo nicety. It is a purchase-deciding, revenue-moving property of AI agents — the difference between €55,000 signed and a nice conversation. If an AI will touch your CRM, your support queue, or your forecast, the question isn’t whether it writes well. It’s whether it finishes what it starts.

Firmulate makes this testable rather than hypothetical. Beyond the benchmark, 242 real, unedited management decisions power a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. (One fairness note: Kimi K3 ran at the API’s default effort setting while the others ran at xhigh, and still nearly won.)

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The dating world figured this out long ago: the person who listens — actually listens, all the way down to the second-layer detail — is the one you commit to. The Firmulate experiment shows AI agents face the same sorting. Every model was honest under pressure. Every model saw the crisis coming. Only the ones that did the reading earned the signature.

So the next time someone — human or machine — impresses you in conversation, ask the Firmulate question: did they read the file? The full results are public at firmulate.com/benchmarks.html, and the company itself runs live for anyone to watch. Chemistry is common. Follow-through is the scarce trait — and now, at least for AI, it has a scoreboard.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

CANCELED | America’s Independence Day parade called off over extreme heat

The July 4th parade in the U.S. has been canceled over dangerously high temperatures, marking a rare cancellation for safety reasons amid ongoing heatwave.

The Future Of AI In Scroll-Driven Depth Technology At Abyssal Station

FABLE/175’s Abyssal Station links scrolling to simulated ocean depth, lighting, pressure, particles and animated marine life.

20-Oz STANLEY Quencher ProTour Flip Straw Tumbler w/ Leakproof Lid $15

The Stanley Quencher ProTour Flip Straw Tumbler with leakproof lid is available for $15, down from its regular price, offering a durable, portable hydration option.

TIL that the son of the man who welcomed the puritans and fed them when they were starving had his head cut off and put on a spike for 20 years at the same location as the first thanksgiving.

New research confirms that Wampanoag sachem Metacomet’s only son was sold into slavery after King Philip’s War, shedding light on his descendants’ history.