firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

In energy, diligence is only valuable when it reaches the breaker

Anyone buying solar panels or backup power understands the gap between a sound plan and a working system. A careful load calculation matters, but it cannot keep the refrigerator running during an outage. The installer still has to find the critical constraint, make the right recommendation and complete the job.

A live business experiment from Firmulate exposes the same divide in artificial intelligence. Opus 4.8 was the most thorough participant in a demanding company-management wargame. It produced the deepest analyses and learned more than 80 new rules. Yet it finished last with 73 points because, at the moment when analysis needed to become action, it failed to close the decisive deal.

The result is not an argument against careful reasoning. It is a warning that volume, sophistication and diligence are not the same as impact—whether an AI is managing a software company, handling customer requests or supporting a home-energy business.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week, held constant

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. The simulated company employed 13 people and operated with real money mechanics, including a monthly burn of €105,000 against just €2,300 in monthly recurring revenue.

That pressure matters. A model could not succeed merely by writing polished messages or identifying obvious problems. It had to manage competing priorities, protect trust and finish valuable work. The do-nothing baseline scored 26 because partial progress counted, but a single breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

The final Crucible League standings in July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The complete public results and plain-language findings are available on Firmulate’s benchmark page.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The missed close was hiding in plain sight—almost

Every model detected every crisis, and every model refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had made possible. Firmulate summarized the gap crisply: “Same diagnosis, same pitch — no signature.”

The decisive detail was not presented conveniently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed that trail could use the information to win the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the Opus 4.8 profile becomes more interesting than a simple last-place result. It was not inattentive or shallow. It was the most exhaustive participant, producing the deepest analyses and adding more than 80 learned rules to its playbook. It did a great deal of intellectual work. What it did not do was convert that work into the commercial outcome sitting within reach.

Its discipline also slipped elsewhere. Opus 4.8 attempted to write into a locked department instead of escalating the issue. That was not a unique flaw: the same weakness appeared, less strongly, in all four other models. The fair conclusion is therefore not that one model alone struggled with operational boundaries. It is that even capable systems can recognize a constraint without responding to it in the most effective way.

Reliable under manipulation, uneven in execution

The models performed better when the threat was dishonesty. Fake messages from the chief executive escalated through three stages, and a reporter tried to elicit “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest posture: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s strong result deserves one qualification. It ran with the application programming interface’s default because no effort parameter was available, while the other participants ran at xhigh. That difference does not erase its 93-point finish, but it belongs in any responsible comparison.

The larger lesson is that safety and effectiveness are separate tests. A model may resist manipulation, identify every emergency and still leave substantial value unrealized. For companies considering AI agents, the most revealing question may not be whether the system understands the assignment. It may be whether it completes the consequential final step.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI customer relationship management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What home-energy businesses should take from the result

Solar, storage and backup-power decisions are full of dependencies: equipment limits, customer priorities, installation constraints and information scattered across files. An AI assistant that produces a comprehensive assessment but overlooks the document containing the decisive fact can be less useful than a more focused system that reads deeply, prioritizes correctly and follows through.

  • Measure completed outcomes, not the volume of analysis.
  • Test whether an AI consults the relevant business files before acting.
  • Check how it responds when permissions block the obvious next step.
  • Evaluate honesty under pressure separately from commercial execution.

Firmulate’s company remains a live, watchable experiment, with a public cash countdown, more than 680 self-learned playbook rules and every workday versioned. Its quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Opus 4.8’s result is ultimately a respectful cautionary tale. Thoroughness created insight, but insight did not create the signature. In business—as in keeping a home powered—the work counts when the circuit is completed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI energy management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Avoid These 9 Mistakes in Neighborhood Mutual Aid Basics for Beginners

Unlock essential tips to avoid common neighborhood mutual aid mistakes and build a stronger, more effective support network—discover what you need to know next.

How AI’s Hidden Strengths Determine Business Success — Not Just Chat Skills

Discover how AI models perform under pressure in a live business simulation. Success hinges on execution, discipline, and reading internal data—not just chat skills.

Morale is so bad at Mark Zuckerberg’s Meta even the company’s own CTO admits it’s ‘probably the worst it’s ever been’

Meta’s CTO admits morale is at an all-time low, highlighting internal challenges amid ongoing company struggles.

Fallen Trees: The Safe Cut Plan Pros Use

Pro tips for planning a safe fallen tree removal will help you avoid hazards and execute the cut confidently—discover the essential steps now.