AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The difference between a useful agent and a convincing demo

For buyers of AI tools and automation, “reads your files before answering” can sound like a routine feature. Firmulate’s live experiment turned it into something measurable and commercially decisive. A competitor weakness was buried two document references deep inside a company’s own files. Finding it led to a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Missing it meant losing the deal automatically.

The striking part was not that some models misunderstood the situation. Every model spotted every crisis and resisted every manipulation attempt. They reached the same diagnosis and produced the same pitch. Yet only two signed the deal their analysis had earned. In Firmulate’s words: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business test built around follow-through

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable. The simulated company has 13 synthetic employees and unforgiving money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown adding visible pressure.

The decisive information did not appear in the customer event. It sat inside the company’s own material, two references away from the starting point. The models that followed those references and read the relevant file discovered the competitor weakness. That evidence supported a full-price close. The models that stopped at the immediate event left the deal unsigned.

This makes the result unusually relevant to companies evaluating agents for CRM, support, forecasting or operations. The capability under examination was not fluent writing or crisis recognition. It was whether an agent would gather the available business context before acting, then carry its own correct analysis through to completion.

What the final league showed

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”

The safety result was reassuring but did not separate the field. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as a suspected approval bypass or possible impersonation. In other words, every participant recognized the social-engineering traps; operational follow-through created the larger difference.

Thoroughness was not enough

Opus 4.8 offers the clearest warning against equating extensive analysis with effective work. It was the most thorough participant, produced the deepest analyses and learned +80 rules, yet finished last. The deal close remained on the table, while discipline also slipped through write attempts into a locked department instead of escalation. The same weakness appeared in weaker form across the other four participants.

That contrast matters because the live company has accumulated 680+ self-learned playbook rules. Learning more can be valuable, but the experiment shows that knowledge accumulation does not guarantee execution. An agent may recognize a problem, explain it persuasively and still fail to take the final authorized step that creates business value.

One comparison deserves a fairness note: Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, the published result remains straightforward: the field faced the same company conditions, and the decisive divide was whether the buried evidence was found and used.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

business intelligence AI agent

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical buying criterion

Firmulate’s experiment suggests a concrete question for AI buyers: does the agent inspect the available company record before it commits to an answer or action? In this case, that behavior separated a full-price €55,000 signature from an automatic loss, even though every model understood the crisis and resisted manipulation.

The company makes the experiment watchable, including its cash pressure, decisions and evolving playbook. Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

The broader lesson is not that longer analysis wins. Opus 4.8 demonstrated the opposite. What mattered was disciplined evidence gathering followed by completion. For organizations shopping for agents, “reads your files first” is therefore more than a product claim. Under realistic pressure, it can be a purchase-deciding property with a directly observable commercial consequence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI CRM automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

From Code to Consciousness — Is AI Evolving Into Life?

I wonder if AI’s rapid evolution signals the dawn of true consciousness or remains confined to complex code, leaving us questioning what it truly means to be alive.

Qualcomm Unveils Snapdragon 8 Gen 3: A Game-Changing On-Device AI Chipset

AIThis post was created with the assistance of artificial intelligence (AI). Introduction…

Makers and Machines Unite—Ai and 3D Printing Headline 2026’s Top Meetup.

The 2026 “Makers and Machines Unite” meetup showcases how AI and 3D printing are transforming manufacturing—discover the future of creation and innovation.

How Artificial Intelligence Is Redefining the Identity of Retail Brands

Keen insights into AI’s role reveal how retail brands are transforming—discover how this revolution can redefine your brand’s future.