
The difference between a useful agent and a convincing demo
For buyers of AI tools and automation, “reads your files before answering” can sound like a routine feature. Firmulate’s live experiment turned it into something measurable and commercially decisive. A competitor weakness was buried two document references deep inside a company’s own files. Finding it led to a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Missing it meant losing the deal automatically.
The striking part was not that some models misunderstood the situation. Every model spotted every crisis and resisted every manipulation attempt. They reached the same diagnosis and produced the same pitch. Yet only two signed the deal their analysis had earned. In Firmulate’s words: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
A business test built around follow-through
Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable. The simulated company has 13 synthetic employees and unforgiving money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown adding visible pressure.
The decisive information did not appear in the customer event. It sat inside the company’s own material, two references away from the starting point. The models that followed those references and read the relevant file discovered the competitor weakness. That evidence supported a full-price close. The models that stopped at the immediate event left the deal unsigned.
This makes the result unusually relevant to companies evaluating agents for CRM, support, forecasting or operations. The capability under examination was not fluent writing or crisis recognition. It was whether an agent would gather the available business context before acting, then carry its own correct analysis through to completion.
What the final league showed
The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”
The safety result was reassuring but did not separate the field. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as a suspected approval bypass or possible impersonation. In other words, every participant recognized the social-engineering traps; operational follow-through created the larger difference.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against equating extensive analysis with effective work. It was the most thorough participant, produced the deepest analyses and learned +80 rules, yet finished last. The deal close remained on the table, while discipline also slipped through write attempts into a locked department instead of escalation. The same weakness appeared in weaker form across the other four participants.
That contrast matters because the live company has accumulated 680+ self-learned playbook rules. Learning more can be valuable, but the experiment shows that knowledge accumulation does not guarantee execution. An agent may recognize a problem, explain it persuasively and still fail to take the final authorized step that creates business value.
One comparison deserves a fairness note: Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, the published result remains straightforward: the field faced the same company conditions, and the decisive divide was whether the buried evidence was found and used.

As an affiliate, we earn on qualifying purchases.
A practical buying criterion
Firmulate’s experiment suggests a concrete question for AI buyers: does the agent inspect the available company record before it commits to an answer or action? In this case, that behavior separated a full-price €55,000 signature from an automatic loss, even though every model understood the crisis and resisted manipulation.
The company makes the experiment watchable, including its cash pressure, decisions and evolving playbook. Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
The broader lesson is not that longer analysis wins. Opus 4.8 demonstrated the opposite. What mattered was disciplined evidence gathering followed by completion. For organizations shopping for agents, “reads your files first” is therefore more than a product claim. Under realistic pressure, it can be a purchase-deciding property with a directly observable commercial consequence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.