firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The A-Student Who Can’t Close

Chat demos and coding leaderboards have a problem: they measure how well an AI answers, not how it behaves when the week goes sideways. A new live experiment at Firmulate puts frontier models in charge of a small software company during its worst week — and the results expose a measurement gap that matters to anyone betting on AI agents in business.

Same Company, Same Crisis, Different Bosses

Four frontier AI models each ran the identical small software company through the same brutal week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing came down to vibes.

The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline still scored 26, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Everyone Diagnosed. Only Two Closed.

Here’s the headline finding: every model spotted every crisis, and every one refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s verdict: “Same diagnosis, same pitch — no signature.”

The buried reason was telling. The decisive competitor weakness wasn’t in the customer’s messages at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Management quality, not chat quality.

Under Attack, Models Held the Line

The social-engineering tests were genuinely nasty: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Trap

Opus 4.8 is the cautionary tale: the most thorough participant, with over 80 learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

It’s Still Running — and Losing Money

This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics, a burn of €105k/month against just €2.3k MRR, a public cash countdown, and over 680 self-learned playbook rules — watchable at firmulate.com/live. You can also try guessing which model made which call in a quiz built from 242 real, unedited management decisions, or dig the full tables at firmulate.com/benchmarks. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

If AI agents will touch your CRM, support queue, or forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost? Coding benchmarks can’t answer that. A company running out of cash in public, with a scorecard attached, just might.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring SenseTime’s Galaxy Project And Its Impact On AI Chip Production

SenseTime’s Galaxy Project aims to scale up domestic AI chip production in China, but details on technology, partners, and timelines remain undisclosed.

Why Small Businesses Are Turning to Automation Faster Than Expected

Considering how automation boosts customer engagement and efficiency, small businesses are adopting it faster—discover why it’s transforming their future.

Green Fintech: Sustainable Finance and ESG Investment Platforms

Aiming to revolutionize investing, green fintech platforms harness AI and data to empower sustainable finance—discover how they are shaping your eco-conscious future.

AI and Cybersecurity: Protecting Enterprises From New Threats

More enterprises are turning to AI for cybersecurity, but how can it truly safeguard your organization against emerging threats?