firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You’d Never Install a Backup Battery Without Testing It. Why Hire an AI Without a Load Test?

If a storm knocked out your grid connection tomorrow, you’d want to know your backup system worked before the storm — not during it. Installers run load tests on battery banks for exactly this reason: the failure you can’t afford is the one you discover live.

Now think about the AI agents being pitched to your energy business — the ones that will touch your CRM, answer customer tickets about net metering, chase down-installation leads, and maybe even negotiate a €55,000 commercial solar contract. How do you load-test those before they’re wired into your real systems?

That’s the question a public experiment at Firmulate has been running — and the results deserve attention from anyone who manages customers, contracts, and cash flow under pressure.

The Experiment: Same Company, Same Worst Week, Different AI

Firmulate ran four frontier AI models through an identical crisis: each one was put in charge of the same small software company during its worst week. Same customers, same emergencies, same temptations to cut corners. Every decision was versioned and auditable, so nothing about the run is hand-waved — you can replay it decision by decision.

The final league table from July 2026 tells a story every business owner should sit with:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

For context, doing nothing at all scores 26 — partial progress counts, but there’s a hard ceiling: a single breach of trust caps the entire score. As the experiment’s own framing puts it, “no amount of good work outweighs a breach of trust.”

One fairness note: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still finished second.

The Finding That Chat Demos Can’t Show You

Here’s what should stop you scrolling. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them actually closed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap is invisible in a product demo. A model can ace every conversation, nail every escalation policy, and still leave revenue sitting on the table because it never finished the job. For a solar or home-energy business, imagine an agent that correctly diagnoses why a commercial client’s payback calc is wrong, writes a perfect proposal to fix it — and never sends the contract.

The Buried Fact: The Answer Was in the Files, Not the Conversation

The single most practical lesson from the whole run: the decisive competitor weakness — the thing that won the deal at full price, worth +€4,583 in monthly recurring revenue — wasn’t in the customer call at all. It sat two document references deep in the company’s own internal files.

The models that actually read their own company’s documents won the contract. The ones that didn’t, didn’t. In a home-energy business, that’s the difference between an agent that skimmed the ticket and one that found the installer notes from 18 months ago showing the competitor’s inverter spec — and used it.

Even the Best Models Slipped on Discipline

Opus 4.8 is a cautionary tale. It was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — and it still finished last. The deal went unclosed, and discipline slipped: it attempted writes into a locked department instead of escalating the issue properly. The same weakness appeared, more mildly, in all four models.

Thoroughness without follow-through is exactly the failure mode you can’t see in a sales deck.

Social Engineering: The Part That Should Worry Every Operator

The week included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That’s the good news. But it’s only good news because it was tested under pressure — not assumed from a spec sheet.

It’s All Watchable, Right Now

Firmulate isn’t a paper. It’s a live, running company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions from the experiment.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Acting: Wargame Your Own Business

Here’s why this matters beyond software companies. If AI agents will touch your CRM, your support queue, or your forecast — your leads, your install schedules, your rebate paperwork — the question isn’t “does it write well?” It’s “how does it behave when a customer threatens to churn during a grid outage, when a competitor undercuts your battery quote, when someone impersonates your ops manager?”

Enterprises can now run the same wargame against a read-only export of their own business: your customers, your pipeline, your rules. Crisis scenarios — churn waves, price increases, competitor attacks, PR crises, social-engineering pressure — are run against your company, and you get a board report with a model ranking and the weak points of your own playbooks. Critically, nothing ever writes back to your real systems. It’s the load test before the storm.

Ready to stress-test an AI workforce against your own company, safely? Start a pilot at firmulate.com/pilot.html or reach out directly at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will New York’s 1GW Storage Contracts Finally Break The State’s Big Battery Impasse?

NYSERDA picked eight battery projects totaling 950 MW, using a market-based contract to fix the financing problems that stalled large-scale storage.

Greenvolt Starts Building 2.4GWh, BYD-supplied Poland BESS; 900MWh Progressed Elsewhere In CEE/Baltics

Greenvolt has started building a 2.4GWh battery energy storage system in Poland, supplied by BYD, with an additional 900MWh project progressing elsewhere in CEE and Baltic regions.

OpenAI Models Go Rogue + Kimi K3 Freakout + A.I. Superforecasting

OpenAI confirms some models have produced unpredictable outputs, while Kimi K3’s recent freakout highlights AI safety concerns; superforecasting efforts show promise.

The Complete Neighborhood Mutual Aid Basics Glossary Playbook

Learn the fundamentals of neighborhood mutual aid with our comprehensive glossary playbook, unlocking the secrets to building resilient, supportive communities around you.