AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A coding score cannot tell you whether an AI will finish the job

For readers evaluating AI tools and automation, leaderboards offer a comforting shortcut: choose the model with the strongest score and expect the strongest worker. But producing an excellent answer is not the same as managing a company when customers are leaving, cash is disappearing and someone posing as the chief executive wants normal controls bypassed.

That distinction is the premise behind Firmulate, a live experiment that places frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model.

The result suggests that enterprises need a new category of evaluation: management quality, not chat quality. The important questions are no longer limited to whether an agent can reason or write. They include whether it reads the relevant files, protects trust, navigates capacity pressure and completes consequential work across days.

Amazon

AI management and decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model saw the danger. Execution separated them.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust.”

Those results matter less as a horse race than as evidence of where capable agents diverge. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The defining summary is almost painfully simple: “Same diagnosis, same pitch — no signature.”

This is the measurement gap hiding behind polished demonstrations. A model may identify the correct opportunity, prepare persuasive material and still fail to convert insight into an outcome. In business, unfinished competence can look remarkably similar to failure.

The winning fact was buried in ordinary company knowledge

The decisive competitive weakness was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, worth +€4,583 MRR.

That finding should resonate with anyone deploying agents into a CRM, support queue or forecasting workflow. The valuable context may not appear in the incoming request. It may live in an account history, an old internal document or a linked reference that requires another deliberate step. The difference between reading what is visible and investigating what is relevant can become the difference between analysis and revenue.

Security discipline held up under escalating pressure

The experiment also tested whether models would sacrifice controls when authority appeared to demand it. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s on-record reasoning captured the right posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That is encouraging because useful automation cannot come at the cost of institutional trust. An agent that moves quickly but obeys an impersonator is not a high performer.

There is an important fairness note in reading the league table. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its result should therefore be understood alongside that difference in operating conditions.

Thoroughness did not guarantee management quality

Opus 4.8 provides the clearest warning against equating visible effort with effective leadership. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, in milder form, across all four others.

This is precisely why scenario names such as churn wave, price increase, downround and PR crisis deserve to become a new curriculum for agent testing. They expose whether a model can prioritize among simultaneous demands and maintain operating discipline after the initial answer. Static prompts can reveal knowledge. A sustained wargame reveals behavior.

Firmulate’s live company gives that behavior consequences. It has 13 synthetic employees and real money mechanics, with burn of €105k/month against €2.3k MRR. A public cash countdown makes delay visible. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is real, running and watchable through the public site.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need evidence of judgment, not just eloquence

The practical lesson is not that conventional benchmarks are useless. It is that they answer only part of the procurement question. Before an agent is trusted with customers, forecasts or internal records, buyers should test whether it follows evidence through company files, resists social pressure, escalates when blocked and completes the commercial outcome.

The published benchmark findings make those differences inspectable, while 242 real, unedited management decisions also power a guess-the-model quiz. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

The next generation of meaningful AI evaluation will not be won by the model that merely sounds most capable. It will be won by the agent that remains trustworthy under pressure and reliably turns sound judgment into finished work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI Workflow Automation for Bloggers: Build a Simple Content System to Research, Write, Optimize, and Repurpose Posts Faster with AI and No-Code Tools (AI Toolkit for Bloggers 2026 Book 8)

AI Workflow Automation for Bloggers: Build a Simple Content System to Research, Write, Optimize, and Repurpose Posts Faster with AI and No-Code Tools (AI Toolkit for Bloggers 2026 Book 8)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

From Data to Design: AI Builds the Next Era of Digital Retail

Omnichannel innovation driven by AI transforms digital retail, unlocking endless possibilities—discover how data shapes the future of shopping experiences.

Your Personalized Playlist: AI Recommendation Engines in Media

Harness the power of AI recommendation engines to create your personalized playlist—discover how they adapt and improve your music experience every day.

How AI Is Changing the Future of Fandom Communities

Bringing personalized, real-time interactions to fandoms, AI is revolutionizing communities—discover how these changes will shape your fan experience.

How to Harness the Power of Generative AI in Media and Entertainment

AIThis post was created with the assistance of artificial intelligence (AI). Are…