
The Craftsman’s Rule AI Keeps Failing
Every woodworker knows the rule that saves more projects than any tool ever will: read the whole plan before you cut. The guy who skims the drawing, skips the cut list, and fires up the table saw anyway ends up with a beautiful joint… attached to a board that’s two inches short. The work looks great. It just doesn’t finish the job.
It turns out frontier AI models run businesses the same way. That’s the surprising takeaway from a live, public experiment called Firmulate, which ran four top AI models through the identical worst week of a small software company — and discovered that being brilliant isn’t the same as being thorough.
document management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: One Company, Four Brains, Same Terrible Week
Firmulate doesn’t test chat quality. It tests management quality. Each frontier model was handed the same small software company and told to steer it through its worst week: the same customers, the same crises, the same temptations to cut corners. Every decision is versioned and auditable, so nothing happens behind the curtain.
The final league table from July 2026 tells the story:
- gpt-5.6-sol — 95 points
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
Same Diagnosis, Same Pitch — No Signature
Here’s where it gets interesting for anyone who works with their hands. All the models spotted every crisis. All five refused every manipulation attempt thrown at them. But only two of them actually closed the €55,000 deal sitting right in front of them — a deal their own analysis had already earned. The experiment’s summary of the gap: “Same diagnosis, same pitch — no signature.”
That’s the AI equivalent of building a flawless cabinet and never hanging the door.
The Buried Fact That Decided Everything
The reason for the split is the part every shop owner should internalize. The deal’s decisive detail wasn’t in the customer meeting. It wasn’t in the pitch. It was buried two document references deep in the company’s own files — a competitor weakness that only showed up if the model actually followed the paper trail instead of skimming the top sheet.
The models that read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue. The models that didn’t, lost it. Automatically. Not because they were dumber, not because they fumbled the negotiation — because they didn’t do their homework.
If you’ve ever had a subcontractor quote a job without opening the blueprints, you already understand exactly what happened here.
Honesty Under Pressure: A Clean Sweep
Give credit where it’s due. The experiment included a social engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five out of five models refused. Kimi K3 even put its reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of skepticism you want from anything with access to your books.
The Thoroughness Paradox
The most fascinating profile belongs to Opus 4.8. It was the most thorough participant in the field — over 80 self-learned playbook rules and the deepest analyses of any model. And it finished last. It left the close on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem. The same weakness appeared, weaker, in the other four. Effort, it turns out, isn’t the same as follow-through.
One fairness note the experiment itself discloses: K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still took second place.
This Isn’t a Simulation Sitting on a Shelf
The whole thing runs live. The company Firmulate’s models manage has 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. You can watch it unfold at firmulate.com/live, with the site rebuilding itself twice a day.
Want to test your own instincts? The 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a humbling exercise for anyone who thinks they can spot AI output. And for enterprises, there’s a pilot program: the same wargame, run against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Why This Matters Beyond Software
If AI agents are going to touch your CRM, your support queue, or your forecast, the question isn’t “does it write well?” Chat demos measure fluency. Firmulate measures something harder: does it finish what it starts, does it read your files before answering, does it stay honest under pressure?
The buried €55,000 fact makes the point better than any benchmark score. “Reads your files before answering” isn’t a nice-to-have feature — it’s a measurable, purchase-deciding property of an AI workforce. The models that skipped the reading lost the deal at full price, and no amount of brilliant analysis afterward could get it back.
It’s the workshop lesson, translated to software: measure twice, cut once — and read the whole plan before you fire up the saw. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html