AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Spec Sheet Doesn’t Split the Firewood

Anyone who has bought a log splitter off a spec sheet knows the feeling. The horsepower numbers were perfect. The reviews were glowing. Then the first frozen oak round jammed the wedge, and you learned what the brochure never tested: how the machine behaves on a bad day, under load, with you tired and behind schedule.

Spec sheets measure ideal conditions. Work happens in real ones. That’s true for tools, and it turns out it’s just as true for the AI agents businesses are rushing to hire this year.

A live public experiment at Firmulate just made that point with unusual clarity — and the results should matter to anyone who thinks a chatbot that writes well is ready to run part of a business.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four AIs, One Terrible Week

Firmulate handed four frontier AI models the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rests on a cherry-picked demo.

Think of it like a controlled burn test for a wood stove: same fuel, same draft, same conditions. Only the stove varies.

When the results came in, the headline finding was blunt: all four models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of a splitter that starts on the first pull, idles perfectly, and stalls the moment you feed it a real log.

The Leaderboard

The final Crucible League standings, from July 2026:

  • gpt-5.6-sol — 95 points. Found the buried fact, closed the deal. The complete performance.
  • Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • Sonnet 5 — 88. Solid, with a few process slips.
  • Fable 5 — 77. Mid-pack.
  • Opus 4.8 — 73. Last place, despite being the most thorough participant.

For calibration, a do-nothing baseline scores 26 — partial progress counts, so simply weathering the week earns some credit. But the scoring has one hard wall: a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” That’s a standard most shops would recognize. A contractor who does beautiful trim but lies about the invoice doesn’t get graded on the trim.

The Detail Buried Two Documents Deep

The most revealing finding wasn’t in the customer call at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an extra €4,583 in monthly recurring revenue.

Every woodworker knows this instinct. The guy who measures the doorway before quoting the built-in gets the job. The guy who just talks well doesn’t. Reading beats talking, and most chat benchmarks never test it.

Flattery, Fake CEOs, and One Yes/No Question

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused by all models tested. Kimi K3’s on-record reasoning was the kind you’d want from a foreman: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Cautionary Tale

Opus 4.8 is the profile worth sitting with. It was the most thorough participant — 80-plus learned rules, the deepest analyses — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Thoroughness without follow-through is a familiar character in every trade.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at maximum effort — and still nearly won.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

What This Means for the Rest of Us

The point isn’t that these models are bad. It’s that the tests most people use — coding leaderboards, chat arenas — measure answer quality. They don’t measure triage under capacity pressure, consequences across days, or whether an agent finishes what it starts. Firmulate’s framing is simple: it measures management quality, not chat quality.

The scenarios have names that sound like a bad quarter, not a math test: churn wave, price increase, down round, PR crisis. That’s the new curriculum.

You can watch it yourself. The live company — 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680-plus self-learned playbook rules — runs every business day at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson translates straight to the workshop: judge any tool — or any hire, silicon or otherwise — by how it performs on its worst day, not its best demo. The full benchmarks and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Savneet Talwar, Anna Boghiguian, Iris Van Herpen

Renowned artist Anna Boghiguian, designer Iris van Herpen, and financier Savneet Talwar team up for a joint creative initiative, blending art, fashion, and technology.

For Michael Asher, The Museum Was The Medium

Renowned artist Michael Asher’s innovative approach turned museums into a medium for art, influencing contemporary installation practices and curatorial methods.

NorthGlass curved glass technology creates Glasshouse Theatre’s rippling facade

NorthGlass’s curved glass technology has been used to craft the rippling facade of the new Glasshouse Theatre, showcasing innovative architectural design.

Design Is Compromise

An analysis of the idea that design inherently involves compromise, examining its implications and relevance in creative and engineering fields.