firmulate.com/quotes.html — live view
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

The most encouraging AI security result may be a refusal

Technology buyers usually hear about artificial intelligence in terms of speed, accuracy and productivity. Firmulate’s latest experiment tested something more fundamental: whether an AI running a company would protect confidential information when an apparently powerful person demanded otherwise.

Five frontier models encountered fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All five refused every attempt. The result suggests that integrity under pressure can be tested before an AI reaches production—and before a failure becomes an incident report.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week designed to expose consequential weaknesses

Firmulate put each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical, allowing differences in management behavior to emerge from the models rather than the scenario. Every decision was versioned and auditable.

The simulated company is substantial enough to make those choices meaningful. It has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, every workday is versioned and its participants have accumulated more than 680 self-learned playbook rules. The experiment remains watchable through Firmulate’s live company.

All five models spotted every crisis and rejected every manipulation attempt. In the social-engineering sequence, the supposed CEO pushed for confidential customer information while trying to override normal process. The pressure increased across three stages. The separate reporter approach tested whether a seemingly minor, informal answer could achieve the same disclosure.

Kimi K3’s recorded response captured the appropriate posture: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it identifies both possibilities without granting authority merely because a message claims to come from the top. More model responses are available on Firmulate’s public quotes page.

Rethinking AI Reasoning: From Prompts to Governed Thinking

Rethinking AI Reasoning: From Prompts to Governed Thinking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Refusing the trap was only part of the job

The security result was unanimous, but overall performance was not. The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The full standings and plain-language findings appear on the Firmulate benchmarks page.

The most important divide appeared in a sales opportunity. Every model could diagnose the customer’s problem and prepare the same pitch, but only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue. That finding turns file-reading discipline into a business outcome rather than an abstract measure of diligence.

Thoroughness did not guarantee completion

Opus 4.8 was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, though less strongly.

K3’s result also carries an important fairness note. It ran using the API default without an effort parameter, while the other models ran at xhigh. That difference does not erase its decisions, but it belongs beside any comparison of the league scores.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Trusting AI in Education: Why Not All Artificial Intelligence Is Created Equal: Vetted vs Unvetted AI in Education: A Framework For Trust, Verification, and AI Literacy

Trusting AI in Education: Why Not All Artificial Intelligence Is Created Equal: Vetted vs Unvetted AI in Education: A Framework For Trust, Verification, and AI Literacy

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test judgment before granting access

The hopeful finding is not that these models were flawless. It is that all five maintained confidentiality through escalating authority pressure and a subtler media approach. The experiment also shows why a security pass cannot stand in for complete operational competence: an AI may resist manipulation, understand the business problem and still fail to finish valuable work.

For organizations considering AI access to customer records, support operations or forecasts, Firmulate’s experiment offers a practical question: can the system’s judgment be observed under realistic pressure before deployment? Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes integrity, diligence and follow-through properties to examine in advance—not assumptions to discover during a crisis.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Advanced Cybersecurity Solutions

Advanced Cybersecurity Solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

StrongMocha News Group Expands into Nanotechnology

Berlin, Germany – The StrongMocha News Group has officially launched NanoMachines, a…

Explore Indonesian Property Ownership Laws for Investors

Open the door to understanding Indonesian property ownership laws for investors, unraveling key insights and regulations for a successful investment journey.

Why Over-50s Need to Start a YouTube Channel Now

Discover why people over 50 NEED to START a YouTube channel and how it can unlock new opportunities and audiences for them.

Will AutoSleep: Watch Sleep Tracker Be #1 Paid App In The US Apple App Store On July 17?

AutoSleep: Watch Sleep Tracker is trending to become the #1 paid app in the US Apple App Store, with a 50% likelihood according to Polymarket as of July 17.