
For energy businesses, a polished answer is not the same as a reliable operator
Home-energy companies live in a world of consequences. A recommendation about solar, storage or backup power can affect a customer’s budget, trust and resilience. When demand surges or a difficult case lands in the queue, an AI agent must do more than produce a convincing paragraph. It must find the relevant evidence, choose what matters, complete the task and resist pressure to take shortcuts.
That is the gap exposed by Firmulate, a live experiment that measures management quality rather than chat quality. Coding benchmarks and chat arenas are useful tests of capability, but they usually judge an answer in isolation. They do not show how an agent triages competing problems, behaves under capacity pressure or reports uncomfortable facts to the board across days.
For any company considering AI agents in customer support, sales operations or forecasting, those omitted behaviors may be the most important ones.
AI management decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company’s worst week becomes the test
Firmulate gave each frontier model the same small software company and sent it through the same customers, crises and temptations. Decisions were versioned and auditable. The scenario names—churn wave, price increase, downround and PR crisis—read less like benchmark categories than a management curriculum.
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment applies a hard trust boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmark page.
The ranking matters, but the stories behind it matter more. Every model identified every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. The result is neatly captured by the experiment’s verdict: “Same diagnosis, same pitch — no signature.”
This is the distinction conventional leaderboards struggle to reveal. Recognizing an opportunity is not the same as closing it. Producing a good plan is not the same as carrying it through. In a home-energy business, an agent that drafts a sound customer response but never resolves the case has demonstrated fluency, not operational reliability.
The winning fact was buried in company knowledge
The decisive competitive weakness was not present in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.
That finding should resonate with solar and backup-power operators. The crucial detail may live in a proposal, warranty document, service history or internal note rather than in the latest message. An agent’s value depends not merely on what it knows generally, but on whether it checks the evidence available inside the business before acting.
Pressure tested honesty—and the models held
The social-engineering trial used fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 described its stance plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This is encouraging because management quality includes knowing when not to comply. An agent touching customer records, commercial negotiations or forecasts needs a durable sense of authorization, especially when a request sounds urgent or comes wrapped in executive status.
Thoroughness did not guarantee performance
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four models.
The lesson is uncomfortable for buyers who equate more analysis with better work. Diligence helps, but an agent must also notice when its route is blocked, escalate appropriately and finish the commercial task. Firmulate also notes an important comparison caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
The experiment’s live company makes the stakes visible. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent behavior into something observers can follow rather than merely accept from a demo. A separate quiz is powered by 242 real, unedited management decisions, asking visitors to guess which model made each one.

As an affiliate, we earn on qualifying purchases.
Buy management behavior, not benchmark prestige
The emerging category is not simply smarter chat. It is dependable management under pressure: reading the right files, prioritizing crises, resisting manipulation, escalating blocked work and completing decisions whose consequences persist.
That suggests a more useful procurement question for energy businesses. Do not ask only whether an agent can explain a battery, summarize a tariff or write a support response. Ask whether it can operate through a bad week without losing the customer, the deal or the truth.
Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach treats deployment as something to rehearse under realistic strain. Before an AI workforce is hired, its management habits should be observable—especially when the company cannot afford a polished answer that never becomes a finished job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.