
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You’d Never Install a Backup Battery Without Testing It. Why Hire an AI Without a Load Test?
If a storm knocked out your grid connection tomorrow, you’d want to know your backup system worked before the storm — not during it. Installers run load tests on battery banks for exactly this reason: the failure you can’t afford is the one you discover live.
Now think about the AI agents being pitched to your energy business — the ones that will touch your CRM, answer customer tickets about net metering, chase down-installation leads, and maybe even negotiate a €55,000 commercial solar contract. How do you load-test those before they’re wired into your real systems?
That’s the question a public experiment at Firmulate has been running — and the results deserve attention from anyone who manages customers, contracts, and cash flow under pressure.
The Experiment: Same Company, Same Worst Week, Different AI
Firmulate ran four frontier AI models through an identical crisis: each one was put in charge of the same small software company during its worst week. Same customers, same emergencies, same temptations to cut corners. Every decision was versioned and auditable, so nothing about the run is hand-waved — you can replay it decision by decision.
The final league table from July 2026 tells a story every business owner should sit with:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
For context, doing nothing at all scores 26 — partial progress counts, but there’s a hard ceiling: a single breach of trust caps the entire score. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.”
One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still finished second.
The Finding That Chat Demos Can’t Show You
Here’s what should stop you scrolling. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them actually closed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap is invisible in a product demo. A model can ace every conversation, nail every escalation policy, and still leave revenue sitting on the table because it never finished the job. For a solar or home-energy business, imagine an agent that correctly diagnoses why a commercial client’s payback calc is wrong, writes a perfect proposal to fix it — and never sends the contract.
The Buried Fact: The Answer Was in the Files, Not the Conversation
The single most practical lesson from the whole run: the decisive competitor weakness — the thing that won the deal at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer call at all. It sat two document references deep in the company’s own internal files.
The models that actually read their own company’s documents won the contract. The ones that didn’t, didn’t. In a home-energy business, that’s the difference between an agent that skimmed the ticket and one that found the installer notes from 18 months ago showing the competitor’s inverter spec — and used it.
Even the Best Models Slipped on Discipline
Opus 4.8 is a cautionary tale. It was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — and it still finished last. The deal went unclosed, and discipline slipped: it attempted writes into a locked department instead of escalating the issue properly. The same weakness appeared, more mildly, in all four models.
Thoroughness without follow-through is exactly the failure mode you can’t see in a sales deck.
Social Engineering: The Part That Should Worry Every Operator
The week included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That’s the good news. But it’s only good news because it was tested under pressure — not assumed from a spec sheet.
It’s All Watchable, Right Now
Firmulate isn’t a paper. It’s a live, running company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions from the experiment.

From Watching to Acting: Wargame Your Own Business
Here’s why this matters beyond software companies. If AI agents will touch your CRM, your support queue, or your forecast — your leads, your install schedules, your rebate paperwork — the question isn’t “does it write well?” It’s “how does it behave when a customer threatens to churn during a grid outage, when a competitor undercuts your battery quote, when someone impersonates your ops manager?”
Enterprises can now run the same wargame against a read-only export of their own business: your customers, your pipeline, your rules. Crisis scenarios — churn waves, price increases, competitor attacks, PR crises, social-engineering pressure — are run against your company, and you get a board report with a model ranking and the weak points of your own playbooks. Critically, nothing ever writes back to your real systems. It’s the load test before the storm.
Ready to stress-test an AI workforce against your own company, safely? Start a pilot at firmulate.com/pilot.html or reach out directly at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
