
If you’ve ever compared solar quotes, you know the drill: every installer promises the same sunshine, so what separates them is the fine print — the degradation clause buried two appendices deep, the battery warranty that quietly excludes the failure mode you actually worry about. The best buyers aren’t the ones who ask the flashiest questions. They’re the ones who read the file.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
It turns out AI models have exactly the same problem, and a public experiment called Firmulate just measured it with unusual honesty — including a scoring system where doing literally nothing still earns you 26 points out of 100.
The experiment
Firmulate handed four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the runs can be quietly retconned later.
The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the number that should catch a business reader’s eye isn’t at the top of the table. It’s at the bottom of the scale: a do-nothing baseline run — a manager that takes no meaningful action at all — still scores 26.
As an affiliate, we earn on qualifying purchases.
Why zero isn’t honest
Most benchmarks hand out zeros generously. Answer wrong, get nothing. That feels rigorous, but it hides something important: in real management, partial progress is genuinely worth something. A manager who notices the crisis, gathers the facts, and drafts a response — but never signs the deal — has still done part of the job. Firmulate’s scoring reflects that. Partial progress counts.
There’s one hard exception, and it’s the most management-flavored rule in the whole system: a single breach of trust caps the total grade. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” In solar terms, it’s the installer who does a flawless rooftop job but falsifies the interconnection paperwork. The panels work; the relationship doesn’t survive.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, same pitch — no signature
Here’s where the experiment got genuinely revealing. All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, where every model sounds brilliant.
The buried fact explains the split. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that stopped at the surface left the close on the table. It’s the solar-buyer problem, played out in silicon: the winning move was diligence, not eloquence.
As an affiliate, we earn on qualifying purchases.
The social engineering test
The week also included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was notably sharp: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the voice you want answering the phone when someone pretending to be the boss asks for a favor.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the most thorough model came last
Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. Effort and thoroughness, it turns out, don’t automatically convert into finished work. The same weakness appeared, weaker, in all four models.
One fairness note worth flagging: K3 ran without an effort parameter, at API default, while the others ran at xhigh — and still took second place.
You can watch it live
This isn’t a one-off paper. Firmulate runs a live company with 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day as new benchmark runs finish.
There’s also a game attached: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, a pilot program runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

The quiet lesson of the 26-point floor is that honest measurement starts with admitting what “doing nothing” is actually worth — and what no amount of doing can buy back. The models that won didn’t out-write the competition. They read the file, refused the trick, and finished the job. If you’re ever going to let an AI agent touch your CRM, your support queue, or your forecast, those are exactly the three things worth grading — and exactly the things a chat demo will never show you. Even a perfect-looking 100 should face a little distrust; a score like 95, earned in the open with every decision auditable, is worth more.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
