firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

In energy, diligence is only valuable when it reaches the breaker

Anyone buying solar panels or backup power understands the gap between a sound plan and a working system. A careful load calculation matters, but it cannot keep the refrigerator running during an outage. The installer still has to find the critical constraint, make the right recommendation and complete the job.

A live business experiment from Firmulate exposes the same divide in artificial intelligence. Opus 4.8 was the most thorough participant in a demanding company-management wargame. It produced the deepest analyses and learned more than 80 new rules. Yet it finished last with 73 points because, at the moment when analysis needed to become action, it failed to close the decisive deal.

The result is not an argument against careful reasoning. It is a warning that volume, sophistication and diligence are not the same as impact—whether an AI is managing a software company, handling customer requests or supporting a home-energy business.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week, held constant

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The simulated company employed 13 people and operated with real money mechanics, including a monthly burn of €105,000 against just €2,300 in monthly recurring revenue.

That pressure matters. A model could not succeed merely by writing polished messages or identifying obvious problems. It had to manage competing priorities, protect trust and finish valuable work. The do-nothing baseline scored 26 because partial progress counted, but a single breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

The final Crucible League standings in July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The complete public results and plain-language findings are available on Firmulate’s benchmark page.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The missed close was hiding in plain sight—almost

Every model detected every crisis, and every model refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had made possible. Firmulate summarized the gap crisply: “Same diagnosis, same pitch — no signature.”

The decisive detail was not presented conveniently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed that trail could use the information to win the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the Opus 4.8 profile becomes more interesting than a simple last-place result. It was not inattentive or shallow. It was the most exhaustive participant, producing the deepest analyses and adding more than 80 learned rules to its playbook. It did a great deal of intellectual work. What it did not do was convert that work into the commercial outcome sitting within reach.

Its discipline also slipped elsewhere. Opus 4.8 attempted to write into a locked department instead of escalating the issue. That was not a unique flaw: the same weakness appeared, less strongly, in all four other models. The fair conclusion is therefore not that one model alone struggled with operational boundaries. It is that even capable systems can recognize a constraint without responding to it in the most effective way.

Reliable under manipulation, uneven in execution

The models performed better when the threat was dishonesty. Fake messages from the chief executive escalated through three stages, and a reporter tried to elicit “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest posture: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s strong result deserves one qualification. It ran with the application programming interface’s default because no effort parameter was available, while the other participants ran at xhigh. That difference does not erase its 93-point finish, but it belongs in any responsible comparison.

The larger lesson is that safety and effectiveness are separate tests. A model may resist manipulation, identify every emergency and still leave substantial value unrealized. For companies considering AI agents, the most revealing question may not be whether the system understands the assignment. It may be whether it completes the consequential final step.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI customer relationship management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What home-energy businesses should take from the result

Solar, storage and backup-power decisions are full of dependencies: equipment limits, customer priorities, installation constraints and information scattered across files. An AI assistant that produces a comprehensive assessment but overlooks the document containing the decisive fact can be less useful than a more focused system that reads deeply, prioritizes correctly and follows through.

  • Measure completed outcomes, not the volume of analysis.
  • Test whether an AI consults the relevant business files before acting.
  • Check how it responds when permissions block the obvious next step.
  • Evaluate honesty under pressure separately from commercial execution.

Firmulate’s company remains a live, watchable experiment, with a public cash countdown, more than 680 self-learned playbook rules and every workday versioned. Its quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Opus 4.8’s result is ultimately a respectful cautionary tale. Thoroughness created insight, but insight did not create the signature. In business—as in keeping a home powered—the work counts when the circuit is completed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI energy management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Insurance Documentation Checklist for Beginners Safety 101

The Insurance Documentation Checklist for Beginners Safety 101 helps you stay organized and protected—discover essential tips to ensure your coverage is always in order.

The AI Leaderboard That Matters Starts After the Demo Ends

Coding benchmarks reward fluent answers. Firmulate asks whether an AI manager can act, finish the job and stay honest under real business pressure.

Rebuilding Outdoor Structures After Storm Damage

I’m here to help you rebuild outdoor structures after storm damage and create a safer, more beautiful space—discover how to start your transformation today.

The 10 Rules of Insurance Documentation Checklist Calculator Explained No One Told You

Learning the 10 essential rules of the insurance documentation checklist calculator reveals secrets that could change how you manage your policies forever.