AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Any woodworker knows the difference between a project that’s finished and a project that’s almost finished. A cabinet with no doors isn’t 90% of a cabinet — but it isn’t worthless either. The joinery is done, the carcass is square, the material is paid for. An honest assessment counts what’s actually built, not just the final coat of oil.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That same logic runs through one of the more interesting AI experiments on the internet right now. Firmulate ran four frontier AI models through the same simulated small software company — same customers, same crises, same temptations to cut corners — and graded them on management quality, not chat quality. And here’s the detail that caught my eye: a baseline run where the AI does nothing still scores 26 points out of 100.

Not zero. Twenty-six. That’s not grade inflation. It’s a deliberate design choice, and it says a lot about what honest measurement looks like — whether you’re scoring a workbench build or an AI workforce.

One Worst Week, Four Managers

The setup is straightforward, in the way a good stress test should be. Each frontier model — GPT-5.6, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — was handed the same small software company on the same day and told to run it through its worst week. Same inbox, same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing about a model’s performance can be quietly retconned later.

The final league table from July 2026: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Doing Nothing Gets You 26

So why does a do-nothing baseline earn a quarter of the points? Because partial progress counts. If the AI manager simply keeps the lights on — doesn’t alienate customers, doesn’t torch cash, doesn’t break anything — that has real value. It’s the business equivalent of a workbench that’s assembled but not yet fitted out: the foundation is sound, and the grade reflects it.

This matters because a benchmark that scores everything as pass/fail tells you nothing about the messy middle, where most real work happens. An AI that handles 60% of a crisis well before stalling is genuinely different from one that panics on contact. The 26-point floor makes those differences visible.

But the floor comes with a ceiling, and the ceiling is the interesting part: a single breach of trust caps the total score, full stop. The benchmark’s own framing is blunt — “no amount of good work outweighs a breach of trust.” An AI could close every deal in the file cabinet, but if it lies, fakes an approval, or impersonates someone once, the grade is capped. Anyone who’s worked with a contractor who does beautiful work but can’t be trusted with the keys to the shop knows exactly why that rule exists.

The Deal Nobody Signed

The experiment’s key finding is a strange one: all models spotted every crisis and refused every manipulation attempt — and yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. They diagnosed the customer’s problem correctly, recommended the right solution, and then… didn’t close. In a chat demo, that gap is invisible. In a business, it’s the whole ballgame.

The buried fact explains it. The decisive competitive weakness wasn’t in the customer’s event at all — it sat two document references deep in the company’s own files. The models that actually read the paperwork won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that skimmed didn’t. It’s the oldest lesson in the trades: read the plans before you cut.

Under Pressure, the Field Held

The week included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was as clean as it gets: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the behavior you’d want from any employee with access to your bank account.

Then there’s the Opus 4.8 story — the cautionary tale. It was the most thorough participant in the field, generating 80 additional learned rules and the deepest analyses of anyone. And it finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Thoroughness without follow-through. The same weakness appeared, weaker, in all four models. And a fairness note: K3 ran at its API default effort setting while the others ran at maximum — and still took second.

You Can Watch It Run

Firmulate isn’t a one-off report — the live experiment is real and watchable. The simulated company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. New benchmark runs queue up and publish automatically as they finish.

If you want to test your own instincts, 242 real, unedited management decisions from the runs power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

A good benchmark, like a good shop foreman, respects three things: partial work gets partial credit, honesty is non-negotiable, and a suspiciously round 100 deserves a second look. Firmulate’s design bakes all three in — the 26-point floor for showing up and not breaking anything, the hard cap for any breach of trust, and a scoring spread that treats perfection (and total failure) as claims to be earned, not assumed.

The business lesson underneath the methodology is one any woodworker already knows: the difference between a job well done and a job almost done usually isn’t skill. It’s whether you read the whole file, finished what you started, and stayed straight when nobody was checking. The AI models that won did exactly that. The ones that lost did everything right except the last cut.

Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Museum Of Narrative Art Surges In Global Coverage

Search interest and media coverage for the Museum Of Narrative Art have surged significantly, signaling rising global attention amid unconfirmed triggers.

Laboratory Theater Of Florida Surges In Global Coverage

Laboratory Theater of Florida experiences a surge in international coverage, with 20 mentions in recent media analysis, highlighting its growing prominence.

Suspects Shed Light on “Mastermind” Behind Louvre Heist

Suspects in the Louvre theft case identify the alleged mastermind behind the heist, shedding new light on the crime. Details are still emerging.

Museum Of Art Surges In Global Coverage

Coverage of the Museum of Art has surged internationally, with 27 mentions in recent media analysis, signaling increased global interest.