AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Automation’s toughest test is not writing—it is finishing

AI tools can draft emails, summarize meetings and produce polished plans. Firmulate is asking a harder question: can an AI workforce operate a company when customers are upset, money is disappearing and shortcuts threaten trust?

Its answer is a live software-business experiment with 13 synthetic employees and real financial pressure. The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its synthetic staff have accumulated more than 680 self-learned playbook rules. Readers can watch the company operate live, turning business automation into an unfolding survival story rather than a controlled product demonstration.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to expose the gap between knowing and doing

Firmulate’s live operation supplies the continuing narrative, while its Crucible League shows what happens when frontier models face the same concentrated test. Each participant ran the same small software company through its worst week, encountering identical customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the benchmark imposed a strict trust boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

All five models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 contract their own work had earned. The result captures a problem familiar to anyone deploying automation: analysis and execution are not the same capability. Firmulate summarizes the failure neatly as “Same diagnosis, same pitch — no signature.”

The decisive clue was already inside the company

The deal did not hinge on eloquence alone. A critical competitor weakness was buried two document references deep in the company’s files rather than presented in the customer event. Models that followed the trail found the fact and won the contract at full price, adding €4,583 in monthly recurring revenue.

That finding gives the experiment unusual relevance for AI-tool buyers. Real organizations rarely arrange all useful context inside one prompt. Important knowledge sits in linked documents, old records and overlooked references. A system may recognize a commercial opportunity and formulate the right pitch, but still fail if it does not inspect the material available to it or complete the final action.

Pressure did not break the trust boundary

The models also faced fake CEO messages that escalated through three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 refused. Kimi K3 described its position on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was a meaningful success. The participants were not merely answering abstract security questions; they were operating amid business urgency and apparent authority. The consistent refusals suggest that frontier models can recognize conspicuous manipulation. The uneven commercial results, however, show that resisting bad instructions does not guarantee effective management.

More thinking did not guarantee a better outcome

Opus 4.8 was the most thorough participant. It produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The comparison also carries an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place result should therefore be read with that difference in mind, not treated as a perfectly controlled measure of model effort.

Firmulate makes the human side of the experiment visible as well. Its public collection of employee statements and decisions lets readers see how the synthetic workforce discusses its situation. That matters because automation failures often look reasonable one message at a time. The larger pattern becomes apparent only when decisions, follow-through and consequences are viewed together.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The most revealing AI demo may be a company that is struggling

Firmulate’s losses are not incidental scenery. The mismatch between €105k in monthly burn and €2.3k in monthly recurring revenue creates a visible test of whether an automated organization can convert awareness into action before its cash runs out.

For businesses evaluating AI tools, the lesson is broader than the league table. A useful agent must locate buried context, protect trust under pressure, respect operational boundaries and finish valuable work. Firmulate’s public experiment shows that today’s models can identify crises and resist manipulation, while still faltering at the moment when analysis must become a completed business outcome.

That tension is what makes the live company compelling: every workday supplies new evidence about whether an AI workforce can learn quickly enough to survive the economics it has been given.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and manipulation detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Politics of AI in Chile Show Why Progress Often Comes With No Easy Victories.

Progress in Chile’s AI politics reveals tough trade-offs, challenging regulations, and security risks that make responsible innovation a complex journey worth exploring.

AI in Fashion and Modeling: Virtual Runways and Digital Models

Here’s how AI in fashion and modeling is revolutionizing virtual runways and digital models, leaving you curious about the future of style and innovation.

The Science Behind How AI Learns What Looks “Good” to You

Beneath AI’s perception of beauty lie complex processes inspired by human vision, revealing fascinating insights that will change how you see aesthetic judgment.

Which Field Dominates 2025: Artificial Intelligence or Data Science?

Beyond doubt, 2025’s tech landscape is dominated by a field that reshapes industries—discover which one is leading and why it matters.