
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Picking an AI agent by its writing is a bet on only part of the job
For teams considering AI tools for customer support, sales or operations, the hard question is whether an agent will act well when the work gets messy: read the relevant records, close a deal and resist pressure to cut corners. Firmulate says its live company experiment puts those decisions on display. Its latest result: Moonshot’s Kimi K3 placed second, ahead of three of four Western frontier models in the field.
AI customer support automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One company, the same difficult week
In the Crucible experiment, each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The company has 13 synthetic employees and real money mechanics; its public cash countdown and workdays continue on Firmulate’s live site.
The July 2026 league table put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The separation came in completing the work: only two signed a €55,000 deal their own analysis had earned. Firmulate describes the gap as “Same diagnosis, same pitch — no signature.”
The file that changed the deal
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That makes the contest a practical test of whether an agent checks available information before acting on a customer-facing opportunity.
K3 found the buried security fact, won the deal, saved the churning customer and resisted all three baits. It did so with one deviation, which Firmulate describes as the cleanest discipline in the field. During a fake CEO message sequence and a reporter’s request for a yes-or-no answer “on background,” all five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says that same weakness appeared, less strongly, in all four models.
A result to inspect, not just a score
Firmulate presents this as a live experiment, not a chat demonstration. The synthetic company burns €105,000 a month against €2,300 in monthly recurring revenue, and its public countdown, employee activity and learned playbook are available to watch at Firmulate. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice. The benchmark page sets out the results and findings in plain language.
For enterprises, Firmulate offers a pilot using a read-only export of a company’s own business. The stated setup does not write back to real systems. That gives buyers a way to examine how an AI workforce handles their own operational context before relying on it.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

AI deal closing and decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the work your AI will actually do
K3’s second-place finish shows that the field is open, while the deal gap shows why model rankings alone do not answer whether an agent will finish a task. If an AI tool may touch your CRM, support queue or forecast, the useful question is how it handles your files, customers and pressure points. Firmulate’s Crucible is one public example of putting those choices under observation before making that bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI ethics and trust management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
