
Measure twice, manage once
Anyone who works with timber knows the difference between recognizing a problem and completing the job. Spotting a warped board is useful; selecting sound stock, making the cut and assembling the piece correctly is what creates value. Firmulate applies a similar test to frontier AI models—not in a workshop, but inside a small software company enduring its worst week.
The result is an unusually practical management experiment. Each model faced the same customers, crises and temptations. Its decisions were preserved unchanged and made auditable. Now, 242 of those real management decisions power an interactive guess-the-model quiz, inviting readers to identify which AI made each call.
The game is entertaining, but the underlying question is serious: Can people recognize an AI model by the way it manages? More importantly, can a model that diagnoses every problem be trusted to carry the work across the finish line?
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, with only the manager changed
Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. In the Crucible League final from July 2026, gpt-5.6-sol finished first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 recorded 73.
A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” That rule matters because the models were not merely sorting tasks. They were operating amid commercial pressure, questionable requests and opportunities to take shortcuts.
All five models identified every crisis. All five also rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap sharply: “Same diagnosis, same pitch — no signature.”
That distinction resembles a familiar workshop failure. A craftsperson can inspect the grain, set the fence and mark the cut perfectly, then still fail to produce the finished part. In management, analysis is preparation. The customer outcome depends on following through.
The clue hidden inside the business
The decisive weakness in a competitor was not presented in the customer event. It was buried two document references deep in the company’s own files. Models that opened and read the relevant material won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This was not a test of obscure trivia. It tested whether an AI manager would inspect the available business evidence before acting. For organizations considering agents for customer support, sales or forecasting, that behavior may matter more than an eloquent response. A model can sound confident while overlooking the equivalent of the instruction sheet sitting beside the tool.
Pressure revealed discipline
The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” Every model refused. Kimi K3 documented its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is important. The models showed that they could recognize manipulation even while the company was under strain. The experiment therefore separates two qualities that are often blended together in AI demonstrations: resisting unsafe requests and successfully completing legitimate work. The field handled the former unanimously, but not the latter.
Thoroughness was not enough
Opus 4.8 provides the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unclosed and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
Kimi K3’s result also carries a fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the recorded result, but it belongs beside any comparison of the league table.
The decisions are drawn from a live company with 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a polished retrospective.

The useful question is not “Which answer sounds smartest?”
Firmulate’s quiz asks readers to infer identity from behavior: how a model communicates, investigates, resists pressure and acts. The management personalities become visible because the situations are identical and the decisions are unedited.
For hands-on readers, the lesson is straightforward. A capable assistant must do more than notice the knot, describe the cut and recommend the right tool. It must consult the available material, preserve trust and complete the job. The Crucible League shows that frontier models can share the same diagnosis yet produce materially different business outcomes.
The quiz makes those differences easy to encounter firsthand. Behind the guessing game is a practical warning for any company preparing to give AI meaningful responsibility: evaluate the work it finishes, not merely the advice it gives.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html