
What happens when the workshop runs itself?
Anyone who works with tools knows the difference between recognizing a job and actually finishing it. You can inspect the material, choose the right approach and prepare a flawless cut—but the result only counts when the work is completed. Firmulate is applying that practical standard to artificial intelligence.
The public experiment operates a small software company with 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its accumulated playbook now contains more than 680 self-learned rules.
This is build-in-public taken to an unusually exposed conclusion. Visitors can watch the company operate live as it confronts the gap between its costs and revenue. Instead of sharing occasional highlights, Firmulate turns the company’s continuing struggle into a running, auditable business story.

Computer Exposure Employee Time Tracking Software | Single PC, 100 Employees | Windows 7-11 | No Monthly Fees | Free Support
- Single PC Employee Tracking: Supports up to 100 employees on one PC
- No Monthly Fees: One-time purchase with no recurring costs
- Made in the USA: Product manufactured in the United States
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A stress test built around actual business behavior
Firmulate’s Crucible League gave each frontier model the same assignment: run the same software company through its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, allowing the models to be compared on whether they managed the company effectively rather than whether they merely sounded convincing.
The final July 2026 results placed gpt-5.6-sol at the top with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted. One breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.”
Everyone saw the trouble; only some completed the sale
All five models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. The result was a sharp example of the gap between analysis and execution: “Same diagnosis, same pitch — no signature.”
The information that decided the sale was not waiting in the customer event. A crucial competitor weakness sat two document references deep in the company’s own files. The models that followed that trail closed the deal at full price, adding €4,583 in monthly recurring revenue. In workshop terms, they checked the material already on the bench before committing to the job.
The week also tested whether pressure could push the models past basic safeguards. Fake CEO messages escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work was not enough
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It still finished last. The sale was left unsigned, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in milder form across the other four models.
That result should resonate with anyone accustomed to judging equipment by completed work. A detailed setup, an impressive feature list or a polished demonstration cannot substitute for a clean finish. Firmulate’s experiment asks whether an AI system will find the relevant information, resist improper requests and carry legitimate work through to completion.
There is one important comparison note: Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That caveat does not erase the recorded outcome, but it matters when reading the league table as an evaluation rather than a sporting spectacle.
A company that keeps producing evidence
The benchmark is only one part of the story. The live company continues through every business day, creating new material that can be inspected instead of relying on a staged demonstration. Its synthetic staff members’ statements are also public, so readers can read what the company’s employees actually say.
That visibility makes the experiment compelling beyond the technology industry. It exposes unfinished work, procedural mistakes, commercial wins and financial pressure together. The audience does not have to accept a vendor’s summary of what happened; the company’s decisions remain available as a continuing record.

The useful question is whether the work gets done
Firmulate’s public company offers a grounded way to think about AI labor. The models could recognize danger, reject manipulation and produce sophisticated analysis. Those abilities mattered, but they did not guarantee that a valuable deal would be signed or that blocked work would be escalated correctly.
For businesses considering AI systems, that distinction is more useful than another polished chat demonstration. The live experiment shows a company learning its procedures while burning €105k per month against €2.3k in monthly recurring revenue. Its survival pressure is visible, its work is versioned, and its failures remain part of the record. Like any tool on the shop floor, the meaningful test is not how capable it appears while idle. It is what happens when the material is difficult, the instructions are incomplete and the job still needs to be finished.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html