AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When diligence becomes a distraction

For businesses evaluating AI tools and automation, polished analysis can be dangerously reassuring. An agent may identify the right problem, resist obvious manipulation and produce an impressively detailed plan—yet still fail at the moment when analysis must become action.

That is the uncomfortable lesson from Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, producing the deepest analyses and adding more than 80 learned rules to its playbook. It also finished last, with a score of 73.

The result is less an indictment of Opus 4.8 than a warning about how organizations assess AI workers. Volume, sophistication and apparent care are not the same as impact. In operational settings, prioritization and follow-through matter just as much as diagnosis.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

Firmulate runs AI models as complete companies and measures their management decisions rather than their ability to answer isolated prompts. In the experiment, each frontier model was asked to operate the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.

The simulated company has 13 employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, while a public cash countdown keeps the consequences visible. Across the broader live company, the agents have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The crucial fact was not in the obvious place

Every model spotted every crisis and refused every manipulation attempt. The sharper divide emerged in a sales opportunity worth €55,000. Only two models signed the deal their own analysis had earned. Firmulate summarized the gap bluntly: “Same diagnosis, same pitch — no signature.”

The deciding information was buried in the company’s own files. A weakness in the competitor’s position sat two document references deep rather than inside the customer event. The models that followed the trail and read the file closed the deal at full price, adding €4,583 in monthly recurring revenue.

That detail should resonate with teams deploying agents across customer relationship management, support or forecasting work. The challenge was not merely recognizing a sales opening. The winning behavior combined investigation, evidence and completion. An agent could understand the situation and still leave the commercial outcome unrealized.

Opus 4.8’s thoroughness was real

Opus 4.8 deserves a fair reading. Its last-place finish did not come from indifference, shallow reasoning or susceptibility to manipulation. It was the most thorough participant, with more than 80 learned rules and the deepest analyses. That diligence is meaningful, particularly in a week designed to create pressure and competing priorities.

It also resisted the experiment’s social-engineering attempts. Fake messages from the CEO escalated over three stages, while a reporter tried to coax out “just one yes/no, on background.” All 5 models refused. Kimi K3 described the threat in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8’s problems were more ordinary—and therefore more instructive. It left the close on the table. Its discipline also slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four models examined for this behavior. In other words, this was not an exotic failure unique to one system. Opus simply displayed the pattern most clearly.

There is also an important qualification around comparative performance. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when interpreting its second-place result.

Management quality is more than intelligence

The contrast between Opus 4.8’s detailed work and its final score exposes a recurring weakness in AI evaluation. A long answer can look more capable than a short one. A growing rulebook can resemble organizational learning. Deep analysis can signal care. None of those qualities guarantees that the agent will identify the decisive task, navigate a blocked path correctly and finish the job.

Firmulate makes that tension tangible through 242 real, unedited management decisions used in its “guess the model” quiz. The decisions invite people to judge behavior without relying on a model’s brand or reputation. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

customer relationship management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The lesson for AI buyers

Opus 4.8’s result is a respectful cautionary tale about mistaking diligence for effectiveness. Its analysis was deep, its learned rules were extensive and its resistance to manipulation held. Yet the deal remained unsigned, and a procedural obstacle prompted the wrong response.

For organizations testing AI agents, the useful question is not simply whether a model can reason or produce comprehensive work. It is whether that reasoning reaches the consequential action while preserving trust and operational discipline. Firmulate’s live experiment shows why those qualities must be observed under pressure: an AI can do the homework, understand the opportunity and still fail to close.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business automation AI solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI sales analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Immerse in Virtual Reality Gaming: Enhancing Experiences With Generative AI

AIThis post was created with the assistance of artificial intelligence (AI). Step…

RHEO: Paint With Light

RHEO is a new app that allows users to create flowing, beautiful light art with simple gestures on iPhone, iPad, and Apple Vision Pro, emphasizing calm and accessibility.

What Makes a Monitor Good for AI Color Work?

Learn what makes a monitor ideal for AI color work and how to choose the perfect display for your professional needs.

12 Ways Generative AI Revolutionizes Media Content Creation

AIThis post was created with the assistance of artificial intelligence (AI). We…