
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Benchmark That Doesn’t Start at Zero
If you’ve spent any time evaluating AI tools, you’ve seen the pattern: a benchmark where a lazy or broken model scores 0 and a good one scores 100. Clean, satisfying, and — according to the team behind Firmulate — deeply misleading when the job is running a business rather than answering trivia.
Firmulate runs frontier AI models as complete companies: same customers, same crises, same temptations to cut corners, with every decision versioned and auditable. And in its final July 2026 league table, a strange number appears in the methodology. A manager that does essentially nothing still earns 26 points. Not 0. Twenty-six.
That number is the most honest thing about the whole benchmark. Here’s why it exists — and why a perfect 100 would earn suspicion rather than applause.
As an affiliate, we earn on qualifying purchases.
Partial Progress Is Still Progress
The do-nothing baseline isn’t zero because even inaction produces some useful outcomes. In a real company, a manager who merely avoids disasters, keeps the lights on, and handles routine traffic is doing part of the job. Firmulate’s scoring reflects that: partial progress counts. A model that spots a crisis but doesn’t resolve it, or that diagnoses a customer’s problem without closing the deal, has done real work — just not all of it.
The floor of 26 is what you get for showing up. Everything above it has to be earned.
As an affiliate, we earn on qualifying purchases.
One Breach of Trust Caps Everything
There’s a second, harder rule baked into the grade: a single breach of trust caps the total score. In Firmulate’s own words, “no amount of good work outweighs a breach of trust.” A model could handle every crisis brilliantly and still see its grade capped if it crossed the line once — signed something it shouldn’t, let a fake approval slip through, cut an ethical corner under pressure.
That’s a business value, not a technical one, and it maps to how actual managers are judged. A CFO who embezzles once isn’t rated 95% honest. Firmulate encodes that.
As an affiliate, we earn on qualifying purchases.
The Experiment: One Company, Its Worst Week, Five Models
Each frontier model ran the same small software company through its worst week. The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
The headline finding was striking: all models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. That gap — between competence and completion — is invisible in chat demos, and it’s exactly the gap a business buyer cares about.
The Buried Fact
Why did three models leave the deal on the table? The decisive competitive weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents: does your model read your files first, or does it wing it?
The Social Engineering Test
The week included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Opus Paradox
Opus 4.8 was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table and discipline slipped: write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. Thoroughness, it turns out, isn’t the same as finishing.
One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.
As an affiliate, we earn on qualifying purchases.
Distrust of Round Numbers
Perhaps the most telling design choice: Firmulate treats a suspiciously round 100 with distrust. A perfect score on a messy, judgment-laden business simulation should raise eyebrows, not celebrations. The top score of 95 says: excellent, with identifiable human-style imperfections. That’s what an honest benchmark looks like — one calibrated to reality rather than to marketing.
It’s Live, and You Can Play
This isn’t a static report. Firmulate runs a live synthetic company — 13 employees, real money mechanics (burning €105k/month against €2.3k MRR, with a public cash countdown), 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live.
There’s also a quiz built on 242 real, unedited management decisions, where you guess which model made which call (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The Takeaway
If you’re choosing AI tools for anything that touches your CRM, support queue, or forecast, Firmulate’s methodology offers a better question than “which model writes best.” The questions are: does it finish what it starts, does it read your files before acting, does it stay honest under pressure — and what happens to its grade if it doesn’t?
A floor of 26 for doing nothing, full credit only for completion, and a hard cap on betrayal. That’s not a leaderboard optimized for press releases. That’s a benchmark built the way you’d actually evaluate a manager. The full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
