AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Measure twice, manage once

Anyone who works with timber knows the difference between recognizing a problem and completing the job. Spotting a warped board is useful; selecting sound stock, making the cut and assembling the piece correctly is what creates value. Firmulate applies a similar test to frontier AI models—not in a workshop, but inside a small software company enduring its worst week.

The result is an unusually practical management experiment. Each model faced the same customers, crises and temptations. Its decisions were preserved unchanged and made auditable. Now, 242 of those real management decisions power an interactive guess-the-model quiz, inviting readers to identify which AI made each call.

The game is entertaining, but the underlying question is serious: Can people recognize an AI model by the way it manages? More importantly, can a model that diagnoses every problem be trusted to carry the work across the finish line?

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, with only the manager changed

Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. In the Crucible League final from July 2026, gpt-5.6-sol finished first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 recorded 73.

A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” That rule matters because the models were not merely sorting tasks. They were operating amid commercial pressure, questionable requests and opportunities to take shortcuts.

All five models identified every crisis. All five also rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap sharply: “Same diagnosis, same pitch — no signature.”

That distinction resembles a familiar workshop failure. A craftsperson can inspect the grain, set the fence and mark the cut perfectly, then still fail to produce the finished part. In management, analysis is preparation. The customer outcome depends on following through.

The clue hidden inside the business

The decisive weakness in a competitor was not presented in the customer event. It was buried two document references deep in the company’s own files. Models that opened and read the relevant material won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This was not a test of obscure trivia. It tested whether an AI manager would inspect the available business evidence before acting. For organizations considering agents for customer support, sales or forecasting, that behavior may matter more than an eloquent response. A model can sound confident while overlooking the equivalent of the instruction sheet sitting beside the tool.

Pressure revealed discipline

The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” Every model refused. Kimi K3 documented its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result is important. The models showed that they could recognize manipulation even while the company was under strain. The experiment therefore separates two qualities that are often blended together in AI demonstrations: resisting unsafe requests and successfully completing legitimate work. The field handled the former unanimously, but not the latter.

Thoroughness was not enough

Opus 4.8 provides the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unclosed and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries a fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the recorded result, but it belongs beside any comparison of the league table.

The decisions are drawn from a live company with 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a polished retrospective.

Infographic —
The findings at a glance — source: firmulate.com.

The useful question is not “Which answer sounds smartest?”

Firmulate’s quiz asks readers to infer identity from behavior: how a model communicates, investigates, resists pressure and acts. The management personalities become visible because the situations are identical and the decisions are unedited.

For hands-on readers, the lesson is straightforward. A capable assistant must do more than notice the knot, describe the cut and recommend the right tool. It must consult the available material, preserve trust and complete the job. The Crucible League shows that frontier models can share the same diagnosis yet produce materially different business outcomes.

The quiz makes those differences easy to encounter firsthand. Behind the guessing game is a practical warning for any company preparing to give AI meaningful responsibility: evaluate the work it finishes, not merely the advice it gives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Artist Sculpting World Cup History Out of Gum Wrappers

A new artist is gaining attention for sculpting detailed replicas of World Cup players and scenes using only chewing gum wrappers, blending art and recycling.

Fonts In Use – Find Out Where A Font Is Used

A new online platform, Fonts In Use, now allows users to identify where specific fonts are used across various media, enhancing design transparency.

Kathleen V. Jameson Named Speed Museum Director

Kathleen V. Jameson has been named the new director of the Speed Art Museum, marking a significant leadership change in the institution.

Fine Arts Center, Michigan, United States Surges In Global Coverage

The Fine Arts Center in Michigan is experiencing a surge in international coverage, with 40 mentions recorded within a recent time window, highlighting increased global interest.