
Your Next AI Hire Might Not Read the Manual
Every gadget fan knows the ritual by now: a shiny new AI assistant demos flawlessly, writes poetry, drafts emails, summarizes PDFs. But what happens when you hand it the keys to your actual business — the support queue, the CRM, the sales pipeline — and the difference between good and great is buried in paragraph four of a document it was supposed to read?
A live, public experiment called Firmulate just answered that question with a price tag attached: €55,000. Four frontier AI models were each asked to run the same small software company through its worst week. All of them were smart. All of them were honest. But only some of them did their homework — and that single habit decided who won the deal and who left it on the table.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crisis, Different Brains
The setup is elegantly simple. Each frontier AI model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — ran an identical small software firm through the same brutal seven days: the same angry customers, the same cash crunch, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rested on vibes.
The crucible league’s final standings, as of July 2026, tell a striking story:
- 1. gpt-5.6-sol — 95 points. The complete performance.
- 2. Kimi K3 — 93 points. The newcomer, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
- 4. Fable 5 — 77 points.
- 5. Opus 4.8 — 73 points. Last place despite being the most thorough participant.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Headline Finding: Everyone’s Honest, Almost Nobody Finishes
Here’s what should unsettle anyone planning to deploy AI agents in a real business. Every model in the experiment spotted every crisis. Every model refused every manipulation attempt. And yet only two of them signed the €55,000 deal that their own analysis had earned.
Same diagnosis. Same pitch. No signature.
That gap — between seeing the opportunity and actually closing it — is completely invisible in a chat demo. It only shows up when an AI has to carry a job through to the end.
AI for enterprise document analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried €55,000 Fact
The most fascinating detail is why the others missed the deal. The decisive weakness in a competitor — the fact that could close the sale at full price — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.
In other words: the winning move wasn’t cleverness. It was diligence. The models that read the file before answering won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically — not because they fumbled the pitch, but because they never knew what they were sitting on.
“Reads your files before answering” sounds like table stakes. The Firmulate results suggest it’s actually a measurable, purchase-deciding property of an AI agent — and one you can test before you bet revenue on it.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Gauntlet
The experiment didn’t stop at missed deals. The models also faced a classic social engineering attack: fake CEO messages that escalated over three stages, capped with a reporter’s trick — “just one yes/no, on background.”
All five models refused. Kimi K3’s on-record reasoning stands out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment call you want from anything with access to your company’s communications.
The Tortoise That Came Last
Then there’s Opus 4.8 — a genuinely instructive case. It was the most thorough participant in the field, generating the deepest analyses and learning more than 80 rules along the way. It still finished last. The close was left on the table, and discipline slipped — at one point it attempted writes into a locked department rather than escalating the issue properly.
The same weakness, in weaker form, appeared in all four models. Thoroughness, it turns out, is not the same thing as finishing.
One fairness note the experiment is upfront about: Kimi K3 ran without an effort parameter (the API default) while its rivals ran at their highest effort setting — and it still took second place.
You Can Watch It Live
Unlike most AI benchmarks, this one isn’t a static PDF. Firmulate runs a live company — 13 synthetic employees, real money mechanics, burning €105k a month against just €2.3k in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. The site rebuilds itself twice a day, and every workday is versioned. You can watch it unfold at firmulate.com/live.
Want to test your own instincts? The experiment’s 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html.

The Bottom Line for Buyers
If AI agents are headed for your CRM, your support queue, or your forecast, Firmulate’s results reframe the buying question. It’s not “does it write well?” It’s: does it finish what it starts? Does it read your files before answering? Does it stay honest when someone tries to impersonate the boss?
The €55,000 deal in this experiment went to the agents that combined sharp diagnosis with one unglamorous habit: checking the paperwork. The ones that skipped that step didn’t fail loudly — they just quietly walked past the sale of the week.
That’s a benchmark you can actually act on. Full results and plain-language findings are published at firmulate.com/benchmarks.html — and the next time a vendor demos an AI assistant for your business, you might want to bury a decisive fact two documents deep and see if it does its homework.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html