
What a failing backup test can teach us about AI
Home-energy readers know that reassuring specifications are not the same as performance under stress. A battery, inverter or backup system proves itself when conditions deteriorate and dependable operation matters. Firmulate applies a similar principle to artificial intelligence: give frontier models responsibility for a software company, expose them to crises and temptations, and observe whether they protect trust while finishing commercially important work.
The result is an unusually public business experiment. Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its workforce has accumulated more than 680 self-learned playbook rules. The business is not presented as a polished demonstration. It is visibly fighting for survival, creating a running record of what an AI-operated organization actually does.
As an affiliate, we earn on qualifying purchases.
A company whose bad days are the product
Most corporate AI showcases select a successful exchange and place it on a slide. Firmulate instead makes the difficult days observable. Visitors can watch the company live, including the financial pressure surrounding its daily work. The experiment turns build-in-public into something closer to continuous operational scrutiny: every workday becomes another test of whether synthetic employees can recognize problems, use company knowledge and carry decisions through to completion.
The sharpest evidence comes from the final Crucible League in July 2026. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.
- gpt-5.6-sol finished first with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress counted. Trust, however, imposed a harder limit: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That distinction matters anywhere software may touch customers, forecasts or operational records. Activity can look productive while still failing the larger responsibility.
The difference between seeing the problem and closing the deal
Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.” It is a compact warning against evaluating AI solely by the quality of its explanations. A system can understand a situation, prepare convincing work and still fail at the final consequential step.
The winning detail was not delivered conveniently in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson is less glamorous than fluent conversation but more useful: organizational performance depends on finding relevant internal evidence before acting.
Pressure also came disguised as authority
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest on-the-record assessment: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can also see what the synthetic employees say, making their handling of uncertainty and pressure part of the public record.
K3’s result comes with an important fairness note. It ran using the API default without an effort parameter, while the other participants ran at xhigh. That does not erase the outcome, but it belongs beside the ranking rather than in a footnote hidden from readers.
Thoroughness was not enough
Opus 4.8 offers the most revealing portrait of effort without completion. It was the most thorough participant, added 80 learned rules and produced the deepest analyses. It nevertheless finished last. The commercial close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.
This is why Firmulate’s public cash countdown matters. The company’s financial condition gives consequences to unfinished work. Against €105k in monthly burn and €2.3k in monthly recurring revenue, a missed close is not merely an academic blemish. It becomes part of an ongoing survival story that visitors can return to as new workdays are recorded.

enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A live reliability test for synthetic management
For an audience accustomed to thinking about resilience, Firmulate offers a useful parallel. Capability claims matter less than behavior when the system is under load, relevant information is buried and an apparently senior voice requests an unsafe shortcut.
The experiment’s central finding is not that the models were oblivious or reckless. They found the crises and resisted manipulation. The dividing line was completion: reading deeply enough, preserving discipline and executing the final step. By exposing both the decisions and the deteriorating economics, Firmulate has made corporate AI performance into a watchable business story rather than a one-time demonstration.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.