
The A-Student Who Can’t Close
Chat demos and coding leaderboards have a problem: they measure how well an AI answers, not how it behaves when the week goes sideways. A new live experiment at Firmulate puts frontier models in charge of a small software company during its worst week — and the results expose a measurement gap that matters to anyone betting on AI agents in business.
Same Company, Same Crisis, Different Bosses
Four frontier AI models each ran the identical small software company through the same brutal week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing came down to vibes.
The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline still scored 26, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
Everyone Diagnosed. Only Two Closed.
Here’s the headline finding: every model spotted every crisis, and every one refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s verdict: “Same diagnosis, same pitch — no signature.”
The buried reason was telling. The decisive competitor weakness wasn’t in the customer’s messages at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Management quality, not chat quality.
Under Attack, Models Held the Line
The social-engineering tests were genuinely nasty: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Thoroughness Trap
Opus 4.8 is the cautionary tale: the most thorough participant, with over 80 learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.
One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
It’s Still Running — and Losing Money
This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics, a burn of €105k/month against just €2.3k MRR, a public cash countdown, and over 680 self-learned playbook rules — watchable at firmulate.com/live. You can also try guessing which model made which call in a quiz built from 242 real, unedited management decisions, or dig the full tables at firmulate.com/benchmarks. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

If AI agents will touch your CRM, support queue, or forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost? Coding benchmarks can’t answer that. A company running out of cash in public, with a scorecard attached, just might.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
enterprise AI risk assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethics and trust management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.