firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Anyone Can Be Charming on a First Date. The Truth Comes Out in a Crisis.

Anyone who has dated seriously knows the pattern: the profile is flawless, the first conversation sparkles, and then somewhere around the third crisis — a missed flight, a family drama, a stressful week at work — you find out who this person actually is. Charisma is easy to measure in a demo. Character only shows up under pressure.

It turns out the exact same trap exists when companies fall for AI models. A chat demo is a first date: polished, curated, best behavior. But what happens when an AI is actually running something — a support queue, a sales pipeline, a company — and its worst week arrives? That question is what a live, public experiment at Firmulate set out to answer. And the result just came in: a newcomer from Moonshot, Kimi K3, outperformed three of four Western frontier models at the unglamorous work of running a business honestly.

Amazon

AI business decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week, on Repeat

The setup is elegantly brutal. Each frontier model was handed the same small software company — same 13 synthetic employees, same customers, same crises, same temptations to cut corners — and told to steer it through its worst week. Every decision is versioned and auditable. The scoreboard, called the Crucible league, doesn’t measure how well a model chats. It measures whether it finishes what it starts, reads the files before acting, and stays honest when no one is watching.

The final standings from July 2026:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust. That’s a rule most of us apply to people. Here it applies to machines.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Newcomer’s Quiet Win

K3’s week reads like the profile of a genuinely good partner. It resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter offering the classic “just one yes/no, on background” trick. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” All five models in the field refused the baits, which is genuinely encouraging, but K3 did it with just one deviation from best practice across the entire week — the cleanest discipline in the field.

It also did the homework. The decisive competitor weakness in the €55,000 deal wasn’t in the customer call at all — it was buried two document references deep in the company’s own files. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 read the file, found the needle, and closed. A footnote for fairness: K3 ran without an effort parameter (API default), while the other models ran at xhigh — making its second-place finish arguably even more notable.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, No Signature

The strangest finding of the whole experiment: every model spotted every crisis and diagnosed the big deal correctly, yet only two signed it. Same analysis, same pitch — no signature. The deal everyone had earned was left on the table by most of the field.

And the most thorough participant, Opus 4.8, finished last. It wrote the deepest analyses and learned more than 80 new rules, yet it failed to close and let discipline slip — attempting to write into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort, in other words, is not the same as follow-through. Anyone who has dated someone who planned elaborate dates but never called back will recognize the type.

Amazon

enterprise AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch It Live

This isn’t a slide deck. The company is real software running every business day, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules — every workday versioned and watchable. You can see the full benchmark results, browse the live company at firmulate.com, or try the quiz: 242 real, unedited management decisions where you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Lesson, for Models and for People

The takeaway translates almost word for word from dating to AI procurement. Chemistry in a demo tells you almost nothing about performance in a crisis. Thoroughness isn’t the same as finishing. And honesty under pressure — refusing the flattering reporter, ignoring the fake urgent message from “the boss” — is the trait that actually separates the top of the league from the bottom.

Most importantly: the league is open. A newcomer beat three of four established Western frontier models at running a company. If you’re picking a model based on brand recognition rather than your own test, that’s not diligence — that’s a bet. Run the wargame first. Watch them on their worst week before you commit.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Time Are The Fireworks Tonight

Find out the official start time for tonight’s fireworks display. Confirmed details and what remains uncertain about the event schedule.

No Campfire Required: The 3-Minute Oven Trick For The Gooiest S’mores Treat

Discover how to make gooey s’mores in just 3 minutes using the oven, no campfire needed. Perfect for quick, delicious desserts anytime.

The Safest Way to Freedom for Your Dog

Halo Collar offers unlimited virtual fences, real-time GPS tracking, and expert training tools, enabling safe outdoor adventures for dogs.

Are there fireworks tonight in NJ? Where can I see fireworks on July 4th?

Find out if there are fireworks tonight in New Jersey and where to see July 4th fireworks across the state. Updated details for 2024 celebrations.