firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A practical question for an automated energy future

For households weighing solar panels, batteries and backup power, impressive promises are not enough. What matters is whether a system notices the detail that changes the decision: a limitation buried in product documentation, a condition attached to a warranty or a fact that becomes important only during an outage.

Businesses adopting AI agents face the same problem. An agent may recognize an urgent situation and produce a polished recommendation, yet still fail because it never opened the document containing the decisive evidence. Firmulate has turned that distinction into something measurable. In its live company experiment, the difference between reading deeply and stopping early determined whether an AI completed a €55,000 sale.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The clue was outside the customer conversation

Firmulate runs AI models as complete companies rather than evaluating isolated chat responses. Each frontier model was asked to manage the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.

The company itself is synthetic but economically unforgiving. It has 13 synthetic employees, burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. That setting makes follow-through consequential: identifying an opportunity is not the same as turning it into revenue.

The most revealing challenge centered on a customer deal. Every model reached the same broad diagnosis and refused every attempt at manipulation. But the fact needed to close the sale was not present in the customer event. It sat two document references deep in the company’s own files, where it revealed a decisive weakness in a competitor’s offering.

Models that found and used that fact secured the agreement at full price, adding €4,583 in monthly recurring revenue. Models that did not read far enough lost the deal automatically. The outcome reduced a fashionable claim—an AI that “reads your files before answering”—to an observable business capability.

Seeing the problem was not enough

All models spotted every crisis. They also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the gap plainly: “Same diagnosis, same pitch — no signature.”

That failure matters because many evaluations reward the visible middle of knowledge work: analysis, drafting and recommendations. Businesses, however, pay for completed outcomes. A system can sound informed while missing a source document, neglecting the final action or failing to carry its own reasoning through to a commercial result.

The final July 2026 Crucible League results put gpt-5.6-sol in the lead with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. Firmulate’s stated principle was uncompromising: “no amount of good work outweighs a breach of trust.”

Thoroughness did not guarantee completion

Opus 4.8 offers the clearest warning against equating visible effort with operational success. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and also attempted writes into a locked department instead of escalating. The same discipline problem appeared in weaker form across all four of the others.

Kimi K3 also requires a fairness qualification. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it finished second and completed the sale. Its handling of social engineering was equally direct.

The models encountered fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.” The result suggests that resistance to obvious pressure may now be common among leading systems, while disciplined research and completion remain stronger differentiators.

Why this matters beyond software sales

In home energy, a recommendation can affect equipment selection, resilience and household spending. The Firmulate experiment did not test solar or battery advice, but its central lesson travels well: the answer can depend on information located somewhere other than the immediate request.

An AI agent working around customer records, support queues or forecasts must connect events with the documents that explain them. It also has to distinguish between authorized work and pressure to bypass controls. Fluency alone demonstrates neither capability.

Firmulate makes the broader experiment watchable as a live company. It also offers a quiz built from 242 real, unedited management decisions, asking readers to guess which model made each choice. For enterprises, the company offers a pilot that runs the same kind of wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buying criterion hiding in plain sight

The €55,000 result reframes AI-agent evaluation. The useful question is not simply whether a model can understand a crisis or write a convincing response. It is whether the model will locate the evidence, respect boundaries and finish the work its analysis has already justified.

That is a measurable property, not a marketing impression. Firmulate’s experiment shows why organizations should test agents inside realistic workflows before granting them responsibility. Sometimes the winning intelligence is not a brilliant answer. It is the discipline to open the next file.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI data analysis tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Report Downed Power Lines and Damaged Equipment

Understanding how to properly report downed power lines and damaged equipment can prevent injuries and save lives, so learn the crucial safety steps now.

Grid Power Returns: The Safe Switch‑Back Procedure Most People Skip

Properly switching back to grid power is crucial for safety and equipment longevity—discover the essential steps most people overlook to avoid costly mistakes.

Tree Hazards and Utility Lines Safety 101

Just knowing the basics of tree hazards near utility lines can prevent accidents and save lives—discover essential safety tips to stay protected.

The AI Company Publishing Its Own Fight for Survival

Firmulate turns a synthetic software company’s daily survival fight into a public test of whether AI can finish hard work and keep trust under pressure.