firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The A-Student Who Can’t Close

Chat demos and coding leaderboards have a problem: they measure how well an AI answers, not how it behaves when the week goes sideways. A new live experiment at Firmulate puts frontier models in charge of a small software company during its worst week — and the results expose a measurement gap that matters to anyone betting on AI agents in business.

Same Company, Same Crisis, Different Bosses

Four frontier AI models each ran the identical small software company through the same brutal week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing came down to vibes.

The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline still scored 26, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Everyone Diagnosed. Only Two Closed.

Here’s the headline finding: every model spotted every crisis, and every one refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s verdict: “Same diagnosis, same pitch — no signature.”

The buried reason was telling. The decisive competitor weakness wasn’t in the customer’s messages at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Management quality, not chat quality.

Under Attack, Models Held the Line

The social-engineering tests were genuinely nasty: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Trap

Opus 4.8 is the cautionary tale: the most thorough participant, with over 80 learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.

One fairness note worth flagging: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

It’s Still Running — and Losing Money

This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics, a burn of €105k/month against just €2.3k MRR, a public cash countdown, and over 680 self-learned playbook rules — watchable at firmulate.com/live. You can also try guessing which model made which call in a quiz built from 242 real, unedited management decisions, or dig the full tables at firmulate.com/benchmarks. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

If AI agents will touch your CRM, support queue, or forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost? Coding benchmarks can’t answer that. A company running out of cash in public, with a scorecard attached, just might.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Thunderbolt Docks Actually Do and Why They Cost So Much

No other device offers such versatile, high-performance connectivity, but understanding why Thunderbolt docks cost so much reveals their true value and capabilities.

Apple Raises Prices on Macs, iPads by $200 or More on Some Models

Apple has raised prices on certain Mac and iPad models by over $200, confirming a significant price adjustment. The move impacts consumers and the market landscape.

What Digital Product Passports Could Mean for Global Trade

A revolutionary shift in global trade awaits, as digital product passports unlock new opportunities—discover how they can transform your business today.

Why SaaS Start‑ups Are Suddenly Launching Hardware—And Winning

Discover why SaaS startups are launching hardware to revolutionize customer engagement and unlock new growth opportunities—find out what’s driving this bold shift.