firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

For energy businesses, a polished answer is not the same as a reliable operator

Home-energy companies live in a world of consequences. A recommendation about solar, storage or backup power can affect a customer’s budget, trust and resilience. When demand surges or a difficult case lands in the queue, an AI agent must do more than produce a convincing paragraph. It must find the relevant evidence, choose what matters, complete the task and resist pressure to take shortcuts.

That is the gap exposed by Firmulate, a live experiment that measures management quality rather than chat quality. Coding benchmarks and chat arenas are useful tests of capability, but they usually judge an answer in isolation. They do not show how an agent triages competing problems, behaves under capacity pressure or reports uncomfortable facts to the board across days.

For any company considering AI agents in customer support, sales operations or forecasting, those omitted behaviors may be the most important ones.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week becomes the test

Firmulate gave each frontier model the same small software company and sent it through the same customers, crises and temptations. Decisions were versioned and auditable. The scenario names—churn wave, price increase, downround and PR crisis—read less like benchmark categories than a management curriculum.

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment applies a hard trust boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmark page.

The ranking matters, but the stories behind it matter more. Every model identified every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. The result is neatly captured by the experiment’s verdict: “Same diagnosis, same pitch — no signature.”

This is the distinction conventional leaderboards struggle to reveal. Recognizing an opportunity is not the same as closing it. Producing a good plan is not the same as carrying it through. In a home-energy business, an agent that drafts a sound customer response but never resolves the case has demonstrated fluency, not operational reliability.

The winning fact was buried in company knowledge

The decisive competitive weakness was not present in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.

That finding should resonate with solar and backup-power operators. The crucial detail may live in a proposal, warranty document, service history or internal note rather than in the latest message. An agent’s value depends not merely on what it knows generally, but on whether it checks the evidence available inside the business before acting.

Pressure tested honesty—and the models held

The social-engineering trial used fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 described its stance plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is encouraging because management quality includes knowing when not to comply. An agent touching customer records, commercial negotiations or forecasts needs a durable sense of authorization, especially when a request sounds urgent or comes wrapped in executive status.

Thoroughness did not guarantee performance

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four models.

The lesson is uncomfortable for buyers who equate more analysis with better work. Diligence helps, but an agent must also notice when its route is blocked, escalate appropriately and finish the commercial task. Firmulate also notes an important comparison caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

The experiment’s live company makes the stakes visible. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent behavior into something observers can follow rather than merely accept from a demo. A separate quiz is powered by 242 real, unedited management decisions, asking visitors to guess which model made each one.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buy management behavior, not benchmark prestige

The emerging category is not simply smarter chat. It is dependable management under pressure: reading the right files, prioritizing crises, resisting manipulation, escalating blocked work and completing decisions whose consequences persist.

That suggests a more useful procurement question for energy businesses. Do not ask only whether an agent can explain a battery, summarize a tariff or write a support response. Ask whether it can operate through a bad week without losing the customer, the deal or the truth.

Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach treats deployment as something to rehearse under realistic strain. Before an AI workforce is hired, its management habits should be observable—especially when the company cannot afford a polished answer that never becomes a finished job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Field Notes: Neighborhood Mutual Aid Basics That Actually Works

Just when you think neighborhood mutual aid can’t succeed, discover key basics that actually work to strengthen community support.

Before You Start: Neighborhood Mutual Aid Basics FAQ Checklist

The essential FAQ checklist to kickstart neighborhood mutual aid effectively and ensure your community’s needs are met from the very beginning.

Cleaning Gutters and Roofs Safely After Storms

Only by following proper safety tips can you effectively and securely clean your gutters and roof after storms.

Tree Hazards and Utility Lines Safety 101

Just knowing the basics of tree hazards near utility lines can prevent accidents and save lives—discover essential safety tips to stay protected.