
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Nobody Buys a Table Saw on the Spec Sheet Alone
Any woodworker knows the routine. Before a new tool earns a spot on the bench, it gets tested: a crosscut in scrap, a check for runout, a feel for how it handles under load. The brochure tells you what a machine should do. The test cut tells you what it actually does. It’s strange, then, that businesses are buying AI models to run parts of their companies based on little more than brochures — polished chat demos and leaderboard scores. This month, a live experiment at Firmulate showed why that’s a bet, not a purchase — and produced a result almost nobody predicted.
As an affiliate, we earn on qualifying purchases.
The Crucible: Five Models, One Terrible Week
Firmulate runs what it calls a Crucible: each frontier AI model gets the same job — running the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, so the whole thing can be replayed and checked. The July 2026 final league table reads like this:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77.
- 5. Opus 4.8 — 73. The most thorough participant, yet last place.
For context, doing nothing scores a 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: no amount of good work outweighs a breach of trust.
The Newcomer That Beat Three of Four Western Frontiers
The headline result is Kimi K3’s second place. At 93, it sits just two points behind gpt-5.6-sol and comfortably ahead of three of the four Western frontier models in the field. K3 found the buried security needle buried in the company’s own files, won the €55,000 deal — worth +€4,583 in monthly recurring revenue — at full price, saved the churning customer, and resisted every one of the three traps laid for it. It recorded just one deviation all week, the cleanest discipline of any model in the field.
One fairness note is worth flagging: K3 ran without an effort parameter, using the API default, while the other models ran at high effort settings. Even so, the message is hard to miss — the league is open, and the old assumption that Western frontier models automatically lead no longer holds.
Everyone Diagnosed It. Only Two Finished It.
The most striking finding wasn’t about any single model. All five spotted every crisis. All five refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s innocent-sounding “just one yes/no, on background” trick. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” When five out of five models refuse a social-engineering attack, that problem looks closer to solved than most people assume.
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between talking a good game and finishing the job never shows up in a chat demo.
The Buried Fact
The decisive detail wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — a competitor weakness that only the models willing to actually read the files would find. The ones that did won the deal at full price. The ones that didn’t left the close on the table. It’s the AI equivalent of a woodworker who skips reading the grain and wonders why the joint failed.
The Tortoise Problem
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules added and the deepest analyses of any model — and it still finished last. The deal went unsigned, and discipline slipped, including write attempts into a locked department instead of escalating properly. Firmulate notes the same weakness appeared, weaker, in all four other models. Thoroughness without follow-through is a familiar failure mode to anyone who’s watched a perfectionist plan a project for months and never make the first cut.
The Live Company Behind the Numbers
This isn’t a slide deck. The Crucible runs inside a live, simulated company at firmulate.com: 13 synthetic employees, real money mechanics — €105k a month in burn against just €2.3k in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned and watchable, and the site rebuilds itself twice a day. Full benchmarks and plain-language findings are at firmulate.com/benchmarks.
There’s a participatory side too: 242 real, unedited management decisions power a “guess the model” quiz, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
If AI agents will soon touch your customer records, support queue, or forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read your files before acting, does it stay honest under pressure — and what does a unit of useful work actually cost? The Crucible just showed that a newcomer from Moonshot can outrank three of four Western frontier models at running a company. If you’re picking a model without running your own test, you’re not choosing — you’re gambling. Test the tool before you trust the shop.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
