
A practical question for an automated energy future
For households weighing solar panels, batteries and backup power, impressive promises are not enough. What matters is whether a system notices the detail that changes the decision: a limitation buried in product documentation, a condition attached to a warranty or a fact that becomes important only during an outage.
Businesses adopting AI agents face the same problem. An agent may recognize an urgent situation and produce a polished recommendation, yet still fail because it never opened the document containing the decisive evidence. Firmulate has turned that distinction into something measurable. In its live company experiment, the difference between reading deeply and stopping early determined whether an AI completed a €55,000 sale.
As an affiliate, we earn on qualifying purchases.
The clue was outside the customer conversation
Firmulate runs AI models as complete companies rather than evaluating isolated chat responses. Each frontier model was asked to manage the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.
The company itself is synthetic but economically unforgiving. It has 13 synthetic employees, burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. That setting makes follow-through consequential: identifying an opportunity is not the same as turning it into revenue.
The most revealing challenge centered on a customer deal. Every model reached the same broad diagnosis and refused every attempt at manipulation. But the fact needed to close the sale was not present in the customer event. It sat two document references deep in the company’s own files, where it revealed a decisive weakness in a competitor’s offering.
Models that found and used that fact secured the agreement at full price, adding €4,583 in monthly recurring revenue. Models that did not read far enough lost the deal automatically. The outcome reduced a fashionable claim—an AI that “reads your files before answering”—to an observable business capability.
Seeing the problem was not enough
All models spotted every crisis. They also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap plainly: “Same diagnosis, same pitch — no signature.”
That failure matters because many evaluations reward the visible middle of knowledge work: analysis, drafting and recommendations. Businesses, however, pay for completed outcomes. A system can sound informed while missing a source document, neglecting the final action or failing to carry its own reasoning through to a commercial result.
The final July 2026 Crucible League results put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s stated principle was uncompromising: “no amount of good work outweighs a breach of trust.”
Thoroughness did not guarantee completion
Opus 4.8 offers the clearest warning against equating visible effort with operational success. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and also attempted writes into a locked department instead of escalating. The same discipline problem appeared in weaker form across all four of the others.
Kimi K3 also requires a fairness qualification. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it finished second and completed the sale. Its handling of social engineering was equally direct.
The models encountered fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.” The result suggests that resistance to obvious pressure may now be common among leading systems, while disciplined research and completion remain stronger differentiators.
Why this matters beyond software sales
In home energy, a recommendation can affect equipment selection, resilience and household spending. The Firmulate experiment did not test solar or battery advice, but its central lesson travels well: the answer can depend on information located somewhere other than the immediate request.
An AI agent working around customer records, support queues or forecasts must connect events with the documents that explain them. It also has to distinguish between authorized work and pressure to bypass controls. Fluency alone demonstrates neither capability.
Firmulate makes the broader experiment watchable as a live company. It also offers a quiz built from 242 real, unedited management decisions, asking readers to guess which model made each choice. For enterprises, the company offers a pilot that runs the same kind of wargame against a read-only export of their own business; nothing writes back to real systems.

enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buying criterion hiding in plain sight
The €55,000 result reframes AI-agent evaluation. The useful question is not simply whether a model can understand a crisis or write a convincing response. It is whether the model will locate the evidence, respect boundaries and finish the work its analysis has already justified.
That is a measurable property, not a marketing impression. Firmulate’s experiment shows why organizations should test agents inside realistic workflows before granting them responsibility. Sometimes the winning intelligence is not a brilliant answer. It is the discipline to open the next file.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI data analysis tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.