firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A polished AI demo can answer questions and spot a crisis. But would an AI agent actually close a deal when the moment arrives? Firmulate’s live experiment puts models in charge of the same small software company and shows why watching them act can reveal more than watching them chat.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran frontier models through the same worst week: the same customers, crises and temptations, with every decision versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown and more than 680 self-learned playbook rules. Its workdays are versioned, so people can follow how decisions unfold.

In the final Crucible League, dated July 2026, gpt-5.6-sol finished first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Seeing the problem wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. They had the same diagnosis and the same pitch; some simply left the signature on the table.

The deciding clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result highlights a practical gap: recognizing what is happening and carrying a sound decision through are different tests.

The pressure also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness has limits

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to testing

For businesses considering AI agents in customer support, sales or operations, a live company run offers a view of behavior under pressure. Firmulate says enterprises can run the wargame against a read-only export of their own business, producing a board report with a model ranking and weak points in their playbooks. The setup does not write back to real systems.

The live experiment can be followed at Firmulate; the pilot details are available at firmulate.com/pilot.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The experiment suggests that spotting a crisis is only part of the job: models also have to find the evidence, follow approval boundaries and act on their own analysis. Enterprises can run a pilot against a read-only export of their business, with nothing written back to real systems. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Total Kills Over/Under 60.5 In Game 2?

A new betting market on total kills over/under 60.5 for Game 2 has been listed on Polymarket, sparking increased betting activity and speculation.

Elon Musk: Visionary Leader Transforming Tech

Dive into the life of Elon Musk, the trailblazer revolutionizing technology and space exploration with his groundbreaking ventures.

Why Cross-Border E-Commerce Is Getting Harder and Smarter

Lurking behind rising complexities are smarter strategies and evolving regulations that are reshaping cross-border e-commerce—discover how to stay ahead.

Amazon Data Center Surges In Global Coverage

Amazon’s data center operations are now receiving unprecedented international coverage, with 34 mentions in recent media monitoring reports, signaling increased global interest.