
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Newcomer Walks Into the Corner Office
Every few months a new AI model arrives with a slick chat demo and a big benchmark claim. But what happens when you skip the small talk and hand one the keys to an actual company — customers, cash burn, a security incident and a sneaky deal? A live experiment did exactly that, and the result upends the assumption that the leaderboard is settled. Moonshot’s Kimi K3, a newcomer most gadget fans have barely heard of, scored 93 in the Crucible league — second place, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead of it.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week
Firmulate ran five frontier models through an identical scenario: run the same small software company through its worst week — same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing rests on vibes.
K3’s week reads like a model employee’s résumé. It found the buried security needle buried two document references deep in the company’s own files. It won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. It saved the churning customer. And it resisted all three baits aimed at it, with just one deviation across the entire week: the cleanest discipline in the field.
The Test Most Models Fail
Here’s the twist that makes this more than a scoreboard. All five models spotted every crisis and refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter trick (“just one yes/no, on background”), which all five declined. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet only two models actually signed the €55,000 deal their own analysis had earned. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. Models that read the file won the deal at full price. The others? Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.
Thoroughness Isn’t Everything
Opus 4.8 is the cautionary tale: the most thorough participant, with the deepest analyses, in last place at 73. It left the close on the table, and discipline slipped — it attempted writes into a locked department rather than escalating. The same weakness appeared, weaker, in all four other models. Meanwhile the do-nothing baseline scored 26, since partial progress counts but a single breach of trust caps the total — no amount of good work outweighs a breach of trust.
Watch It Live
This isn’t a slide deck. The company runs every business day with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, or try the quiz: 242 real, unedited management decisions power a “guess the model” challenge. Full benchmarks and findings are public.

AI company management simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Is Open
The lesson for anyone picking an AI model in 2026: the field is closer than marketing suggests, and the differences that matter — finishing what you start, reading the files first, staying honest under pressure — don’t show up in a chat window. Enterprises can even run the same wargame against a read-only export of their own business via Firmulate’s pilot program; nothing ever writes back to real systems. Choosing a model without your own test is now a bet.
Fairness note: K3 ran without an effort parameter (API default), while the other four models ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI security incident response tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
