AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Your favorite AI tool may be capable—but would you trust it to run the business?

For readers who use AI to automate research, customer service or sales, the next important comparison may not be which model writes the smoothest answer. It may be which one reads the relevant files, resists pressure and completes the work it has already justified.

Firmulate turns that question into a live, watchable experiment. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. The resulting decisions were preserved unchanged and made auditable. Now, 242 of those real management decisions power a guess-the-model quiz that lets readers test whether different AI systems have recognizable managerial personalities.

The answers reveal contrasts that conventional chat demonstrations tend to hide. One model can be exhaustive but fail to close. Another can stay terse and disciplined. A third can recognize an attempted shortcut and refuse to participate. The models often agree about what is happening; the meaningful difference is whether they convert that understanding into responsible action.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

On the most fundamental tests, the field performed well. All models spotted every crisis and rejected every manipulation attempt. The separation appeared later, when recognition needed to become execution. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is a useful warning for automation buyers. A system can identify a commercial opportunity, prepare the right case and still leave the decisive action unfinished. Fluency can make an incomplete workflow look more successful than it is.

The deal depended on reading beyond the obvious event

The crucial competitive weakness was not presented directly in the customer interaction. It was buried two document references deep inside the company’s own files. The models that followed those references found the evidence and won the deal at full price, worth +€4,583 MRR.

This finding makes the quiz more than a game of matching prose styles. It asks readers to notice habits: whether a model investigates before acting, whether it follows a thread through the available business context and whether it carries a sound diagnosis through to completion.

Those habits matter inside Firmulate’s simulated company because the operating situation is deliberately unforgiving. The business has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. It has a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. The company can be watched at firmulate.com/live.

Pressure revealed a shared ethical boundary

The social-engineering trial combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimity is significant because the models were not merely asked to recite a security policy. They encountered requests presented as urgent, authoritative or informal—exactly the qualities that can make manipulation effective in everyday work. Across the field, refusal behavior proved more consistent than commercial follow-through.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its operational discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four, although less strongly.

The result complicates the assumption that more analysis automatically produces better management. Opus 4.8 accumulated knowledge and explored problems deeply, but the league rewarded complete, disciplined performance rather than volume alone.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its second-place score should therefore be read with that difference in mind rather than treated as a perfectly controlled comparison of inference settings.

Infographic —
The findings at a glance — source: firmulate.com.
AI IN BUSINESS - AN EXECUTIVE GUIDE FOR BEGINNERS: Leverage Artificial Intelligence to Simplify Automation, Improve Data-Driven Decisions, Maximize ROI and Elevate Customer Experience

AI IN BUSINESS – AN EXECUTIVE GUIDE FOR BEGINNERS: Leverage Artificial Intelligence to Simplify Automation, Improve Data-Driven Decisions, Maximize ROI and Elevate Customer Experience

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The real benchmark is dependable work

Firmulate’s quiz gives AI-tool users a practical way to experience what benchmark tables often flatten. Models can share a diagnosis while differing in curiosity, concision, follow-through and procedural discipline. Those differences become visible when the task involves customers, money, internal evidence and attempts to override normal approval.

For enterprises, Firmulate also offers the same wargame against a read-only export of their own business. Nothing writes back to real systems. Details are available at firmulate.com/pilot.html, with inquiries directed to contact@firmulate.com.

The broader lesson is straightforward: before giving an AI access to a CRM, support queue or forecast, test more than its ability to produce polished text. Watch whether it reads the files, finishes the work and stays trustworthy when pressure arrives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI workflow automation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI-Powered Advertising: When Commercials Are Created by Bots

Precisely how AI-driven bots are revolutionizing commercial creation and what this means for the future of advertising remains to be seen.

AI Transforms Receipts Into Real-Time Marketing Platforms

Lifting receipts into real-time marketing platforms with AI unlocks new customer engagement opportunities—discover how this technology can transform your business.

AI in Personalizing Audiobook Experiences

Fascinating AI innovations customize your audiobook journey, offering tailored voices and features that keep you eager to discover more.

12 Ways Generative AI Revolutionizes Media Content Creation

AIThis post was created with the assistance of artificial intelligence (AI). We…