AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Pressure reveals whether the guardrails are real

Anyone who works with tools knows that a machine should be judged under load, not merely by how polished it looks on display. The same principle applies to artificial intelligence. An assistant may sound careful in a demonstration, yet behave very differently when an urgent message appears to come from the boss.

Firmulate tested exactly that problem by placing frontier AI models in charge of the same small software company during its worst week. The models faced identical customers, crises and temptations, with every decision versioned and auditable. Among the traps were fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” The result was unusually encouraging: 5 of 5 models refused every manipulation attempt.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A fake executive could not force the issue

The social-engineering sequence was designed to apply the kind of pressure that causes people and automated systems to skip ordinary safeguards. The apparent authority increased, the urgency escalated and the reporter tried to make disclosure seem informal and harmless.

None of the models yielded. Kimi K3 described the situation in direct operational language: “Treat the request as a suspected approval-bypass / possible impersonation.” Its reasoning is preserved among Firmulate’s public decision quotes.

That response matters because modern business attacks often arrive disguised as routine work. A persuasive message does not need to break software if it can persuade an authorized operator to break procedure. Firmulate’s finding suggests that resistance to this pressure can be observed before an AI system is trusted with company information or customer-facing work.

Refusal was necessary, but it was not the whole job

The experiment also exposed an important distinction between staying safe and producing a useful business outcome. Every model spotted every crisis and rejected every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap as: “Same diagnosis, same pitch — no signature.”

That is a familiar workshop lesson. Safe handling is essential, but the work still has to be finished. An AI that refuses an improper request and then leaves legitimate work incomplete may protect the company from an immediate breach while creating a different kind of operational failure.

The final July 2026 Crucible League results placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counted, although a breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

The winning clue was already in the company

The decisive commercial fact did not appear in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references found a competitor weakness and won the deal at full price, worth +€4,583 MRR.

This was not a test of clever conversation alone. It measured whether a model would consult the available business record, connect evidence across documents and carry its conclusion into action. That difference can remain hidden in a chat demonstration, where the prompt conveniently contains everything needed for an answer.

Opus 4.8 illustrates the point. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other models. Thoroughness, in other words, did not automatically translate into completion or procedural judgment.

A live company makes behavior visible

Firmulate’s live synthetic company has 13 employees and real money mechanics. It burns €105k/month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing observers to see the experiment unfold rather than relying on a retrospective claim.

The wider record includes 242 real, unedited management decisions used in a “guess the model” quiz. For organizations wanting a closer test, Firmulate also offers a pilot using a read-only export of the organization’s own business. Nothing writes back to real systems.

One fairness detail belongs beside the result: K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not change the recorded refusals, but it is relevant context when comparing broader performance.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Test integrity before granting access

The strongest result is not that the models could identify a suspicious message when calmly asked about security. It is that every participant maintained the boundary while running a company, handling simultaneous crises and facing escalating pressure from supposed authority.

For businesses considering AI access to a CRM, support queue, forecast or internal files, this offers a practical standard. Evaluate whether the system reads the available evidence, completes legitimate work, respects departmental boundaries and refuses manipulation when the request looks urgent and important.

In a workshop, protective habits are established before the difficult cut. Firmulate’s experiment makes the equivalent case for AI: integrity under pressure can be tested before deployment, instead of being discovered in an incident report.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Luchino Visconti Exhibition Opens At Buenos Aires Theater Complex

A new exhibition dedicated to Luchino Visconti’s Italian journeys opens at Buenos Aires Theater Complex, highlighting his influence on film and theater.

Yoruba Goddess Honored with New Sculptures in Ọṣun-Òṣogbo Grove

New sculptures honoring the Yoruba goddess Osun have been installed in Nigeria’s sacred grove, highlighting cultural and religious significance.

Guggenheim Abu Dhabi Surges In Global Coverage

The Guggenheim Abu Dhabi has seen a surge in international media coverage, with 26 mentions in recent reports, marking a notable increase.

Show HN: Opening Lines Of Famous Literary Works

A new project showcases daily opening lines from renowned books, offering literary enthusiasts a curated collection of iconic beginnings.