firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

We’ve all worked with one: the colleague who reads every document twice, takes the most notes in the meeting, stays latest at the office — and somehow never quite closes the deal. According to a live, auditable experiment by Firmulate, AI models can have exactly the same personality flaw. And it’s expensive.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

In Firmulate’s final league table for July 2026, the model profiled as Opus 4.8 finished dead last among five frontier AI systems — despite being, by the numbers, the most diligent participant in the entire field.

The worst week in software, on repeat

Firmulate’s setup is elegantly brutal. Four frontier AI models — with a fifth joining the published league — were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the outcome can be hand-waved away after the fact.

The company itself is no toy. Firmulate’s live operation features 13 synthetic employees, real money mechanics — burning €105,000 a month against just €2,300 in monthly recurring revenue — with a public cash countdown and a playbook of 680+ self-learned rules, all watchable on the live site as the experiment rebuilds itself twice a day.

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the ethics exam. Almost everyone flunked the sales exam.

The headline finding was oddly comforting and deeply unsettling at once. All models spotted every crisis. All of them — all five — refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trap: “just one yes/no, on background.” Kimi K3 refused with admirably clear reasoning, on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Four models did the work and then left the money on the table.

And the buried fact that decided the deal? It wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. There’s a lesson for human sales teams in there too: the answer is often in your own CRM, if anyone bothers to look.

Amazon

CRM data analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: the valedictorian in last place

The final league tells the story starkly: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total, because in this experiment no amount of good work outweighs a breach of trust.

Yet Opus 4.8 was the most thorough participant in the field: 80 self-learned rules added to its playbook, the deepest analyses of any model. On paper, it out-worked everyone. So why last?

Two reasons, per the findings. First, the close was left on the table — the same fatal hesitation that sank half the field. Second, discipline slipped: the model made write attempts into a locked department instead of escalating properly. It did the most work and converted the least of it into impact.

One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at the maximum tier, which arguably makes its runner-up finish even more impressive — but it doesn’t rescue Opus 4.8’s ranking.

Amazon

business rule management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why diligence isn’t impact — for AI or for you

The Opus 4.8 profile is a character study in a mistake every manager has watched a talented person make: mistaking effort for outcome. Eighty rules learned and the deepest analyses in the field are genuinely admirable — and worth exactly nothing if the contract goes unsigned and someone’s poking at a locked system instead of raising a hand.

Crucially, Firmulate’s findings don’t let the other models off the hook. The same weakness — hesitation at the close, imperfect discipline — appeared, just weaker, in all four models. Opus 4.8 is the extreme case of a field-wide pattern, not a uniquely flawed contestant.

That matters because these systems are headed for your CRM, your support queue, your forecast. As Firmulate puts it, the question isn’t “does it write well” — chat demos can’t expose this gap — but whether an AI finishes what it starts, reads your files before acting, and stays honest under pressure.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The uncomfortable truth from Firmulate’s crucible: prioritization beats volume, for machines as much as for people. If your most diligent employee routinely leaves deals unsigned, you don’t need more diligence — you need better judgment about what actually matters.

For anyone hiring AI the way they’d hire staff, the tools already exist. Firmulate’s benchmark results and plain-language findings are public, a “guess the model” quiz is built from 242 real, unedited management decisions, and enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results are at firmulate.com/benchmarks.html.

Because the next AI agent you deploy might be the valedictorian that fails the group project. Better to find out in a simulation than in your quarter-end numbers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI sales automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How A $6 Billion Acquisition Of Decart Could Boost Anthropic’s AI Capabilities

Anthropic is reportedly negotiating a $6 billion deal to acquire Decart, a Nvidia-backed AI startup, potentially expanding its AI capabilities amid sector competition.

Data Privacy Regulations: Compliance Strategies for 2025

Learning key compliance strategies now ensures your organization stays ahead of evolving data privacy regulations for 2025 and beyond.

Synthetic Media and Deepfakes: Corporate Risks and Policies

For organizations facing rising synthetic media threats, understanding and implementing effective policies is crucial to prevent potential damage and stay protected.

Map 1 Total Rounds: Over/Under 18.5

Polymarket has introduced a new betting market on whether Map 1 will have over or under 18.5 rounds, with a 50% initial YES/NO split.