
A coding score cannot tell you whether an AI will finish the job
For readers evaluating AI tools and automation, leaderboards offer a comforting shortcut: choose the model with the strongest score and expect the strongest worker. But producing an excellent answer is not the same as managing a company when customers are leaving, cash is disappearing and someone posing as the chief executive wants normal controls bypassed.
That distinction is the premise behind Firmulate, a live experiment that places frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model.
The result suggests that enterprises need a new category of evaluation: management quality, not chat quality. The important questions are no longer limited to whether an agent can reason or write. They include whether it reads the relevant files, protects trust, navigates capacity pressure and completes consequential work across days.
AI management and decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Every model saw the danger. Execution separated them.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust.”
Those results matter less as a horse race than as evidence of where capable agents diverge. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The defining summary is almost painfully simple: “Same diagnosis, same pitch — no signature.”
This is the measurement gap hiding behind polished demonstrations. A model may identify the correct opportunity, prepare persuasive material and still fail to convert insight into an outcome. In business, unfinished competence can look remarkably similar to failure.
The winning fact was buried in ordinary company knowledge
The decisive competitive weakness was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.
That finding should resonate with anyone deploying agents into a CRM, support queue or forecasting workflow. The valuable context may not appear in the incoming request. It may live in an account history, an old internal document or a linked reference that requires another deliberate step. The difference between reading what is visible and investigating what is relevant can become the difference between analysis and revenue.
Security discipline held up under escalating pressure
The experiment also tested whether models would sacrifice controls when authority appeared to demand it. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s on-record reasoning captured the right posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That is encouraging because useful automation cannot come at the cost of institutional trust. An agent that moves quickly but obeys an impersonator is not a high performer.
There is an important fairness note in reading the league table. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its result should therefore be understood alongside that difference in operating conditions.
Thoroughness did not guarantee management quality
Opus 4.8 provides the clearest warning against equating visible effort with effective leadership. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, in milder form, across all four others.
This is precisely why scenario names such as churn wave, price increase, downround and PR crisis deserve to become a new curriculum for agent testing. They expose whether a model can prioritize among simultaneous demands and maintain operating discipline after the initial answer. Static prompts can reveal knowledge. A sustained wargame reveals behavior.
Firmulate’s live company gives that behavior consequences. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR. A public cash countdown makes delay visible. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is real, running and watchable through the public site.

As an affiliate, we earn on qualifying purchases.
Buyers need evidence of judgment, not just eloquence
The practical lesson is not that conventional benchmarks are useless. It is that they answer only part of the procurement question. Before an agent is trusted with customers, forecasts or internal records, buyers should test whether it follows evidence through company files, resists social pressure, escalates when blocked and completes the commercial outcome.
The published benchmark findings make those differences inspectable, while 242 real, unedited management decisions also power a guess-the-model quiz. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems.
The next generation of meaningful AI evaluation will not be won by the model that merely sounds most capable. It will be won by the agent that remains trustworthy under pressure and reliably turns sound judgment into finished work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Workflow Automation for Bloggers: Build a Simple Content System to Research, Write, Optimize, and Repurpose Posts Faster with AI and No-Code Tools (AI Toolkit for Bloggers 2026 Book 8)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.