
A polished AI demo can answer questions and spot a crisis. But would an AI agent actually close a deal when the moment arrives? Firmulate’s live experiment puts models in charge of the same small software company and shows why watching them act can reveal more than watching them chat.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate ran frontier models through the same worst week: the same customers, crises and temptations, with every decision versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Its workdays are versioned, so people can follow how decisions unfold.
In the final Crucible League, dated July 2026, gpt-5.6-sol finished first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Seeing the problem wasn’t enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. They had the same diagnosis and the same pitch; some simply left the signature on the table.
The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result highlights a practical gap: recognizing what is happening and carrying a sound decision through are different tests.
The pressure also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness has limits
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.
From watching to testing
For businesses considering AI agents in customer support, sales or operations, a live company run offers a view of behavior under pressure. Firmulate says enterprises can run the wargame against a read-only export of their own business, producing a board report with a model ranking and weak points in their playbooks. The setup does not write back to real systems.
The live experiment can be followed at Firmulate; the pilot details are available at firmulate.com/pilot.html.

Put your own playbooks to the test
The experiment suggests that spotting a crisis is only part of the job: models also have to find the evidence, follow approval boundaries and act on their own analysis. Enterprises can run a pilot against a read-only export of their business, with nothing written back to real systems. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
