AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Picking an AI agent by its writing is a bet on only part of the job

For teams considering AI tools for customer support, sales or operations, the hard question is whether an agent will act well when the work gets messy: read the relevant records, close a deal and resist pressure to cut corners. Firmulate says its live company experiment puts those decisions on display. Its latest result: Moonshot’s Kimi K3 placed second, ahead of three of four Western frontier models in the field.

Amazon

AI customer support automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, the same difficult week

In the Crucible experiment, each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The company has 13 synthetic employees and real money mechanics; its public cash countdown and workdays continue on Firmulate’s live site.

The July 2026 league table put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The striking result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The separation came in completing the work: only two signed a €55,000 deal their own analysis had earned. Firmulate describes the gap as “Same diagnosis, same pitch — no signature.”

The file that changed the deal

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That makes the contest a practical test of whether an agent checks available information before acting on a customer-facing opportunity.

K3 found the buried security fact, won the deal, saved the churning customer and resisted all three baits. It did so with one deviation, which Firmulate describes as the cleanest discipline in the field. During a fake CEO message sequence and a reporter’s request for a yes-or-no answer “on background,” all five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says that same weakness appeared, less strongly, in all four models.

A result to inspect, not just a score

Firmulate presents this as a live experiment, not a chat demonstration. The synthetic company burns €105,000 a month against €2,300 in monthly recurring revenue, and its public countdown, employee activity and learned playbook are available to watch at Firmulate. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice. The benchmark page sets out the results and findings in plain language.

For enterprises, Firmulate offers a pilot using a read-only export of a company’s own business. The stated setup does not write back to real systems. That gives buyers a way to examine how an AI workforce handles their own operational context before relying on it.

Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI deal closing and decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work your AI will actually do

K3’s second-place finish shows that the field is open, while the deal gap shows why model rankings alone do not answer whether an agent will finish a task. If an AI tool may touch your CRM, support queue or forecast, the useful question is how it handles your files, customers and pressure points. Firmulate’s Crucible is one public example of putting those choices under observation before making that bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why AI Dubbing Could Go Mainstream Faster Than Expected

Linguistic adaptability and natural voice quality are propelling AI dubbing toward mainstream adoption, transforming global content reach—discover how inside.

How AI Is Transforming Podcast Editing and Production

Inevitably, AI is revolutionizing podcast editing and production, offering exciting innovations that can enhance your shows—discover how inside.

Discover The AI Innovation Powering Station 36’S Shortwave Listening Website

Discover how AI-driven web design recreates a vintage shortwave radio room with interactive controls, spectral visualization, and authentic sounds.

Mastering Film and TV Production With AI: a Futuristic Guide

AIThis post was created with the assistance of artificial intelligence (AI). We…