
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A benchmark that treats inaction as evidence
Technology benchmarks usually reward clean, decisive outcomes. That makes a score of zero for doing nothing feel intuitive—and a score of 26 feel suspicious. Firmulate’s Crucible League deliberately challenges that intuition.
Its do-nothing baseline receives 26 points because the test recognizes partial progress. A model can notice trouble, understand a customer, or prepare useful work without completing the decisive action. Those achievements matter, even when the business result never arrives. At the same time, the benchmark draws a hard boundary around trust: a single breach caps the total grade because, in Firmulate’s words, “no amount of good work outweighs a breach of trust.”
That combination makes the Firmulate benchmark unusually revealing. It does not pretend that unfinished work has no value, but it also refuses to let productivity compensate for misconduct. For readers accustomed to polished AI demos, that is a more realistic way to ask whether an agent is ready for responsibility.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The difference between understanding and finishing
Firmulate gave frontier models the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73.
The more important result sits beneath that table. All models identified every crisis and rejected every manipulation attempt, yet only two signed the €55,000 deal that their own analysis had earned. “Same diagnosis, same pitch — no signature” is the experiment’s sharpest summary.
This is why partial credit matters. A model that diagnoses a problem correctly has demonstrated something valuable. A model that develops the right commercial pitch has progressed further. But neither achievement is identical to closing the deal. A benchmark that records only the final signature would erase meaningful differences between doing nothing, investigating well and nearly finishing. One that rewards analysis as if it were execution would conceal the opposite problem.
The do-nothing score of 26 provides a reference point. It shows how much of the evaluation can be satisfied without active management and prevents readers from mistaking any positive score for genuine performance. The useful comparison is therefore not merely whether a model scored above zero, but how far it moved beyond inaction—and whether it converted its work into a business outcome.
The decisive fact was hiding in plain sight
The deal also tested whether models would use the company’s own knowledge. The competitor weakness that mattered was buried two document references deep in internal files, rather than presented in the customer event. Models that followed the references found it and won the deal at full price, worth +€4,583 MRR.
For businesses considering AI agents, this distinction is crucial. A system may respond intelligently to the message in front of it while missing the evidence already available elsewhere in the organization. Firmulate’s result turns file-reading discipline into a visible business consequence: full-price revenue rather than an impressive but incomplete analysis.
Trust is a ceiling, not a bonus
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
Firmulate’s trust cap reflects the asymmetry of real managerial work. Useful actions accumulate gradually, but a serious violation can overwhelm them. That is especially relevant for agents operating around customer records, financial forecasts or support conversations. A benchmark can acknowledge competent work while still refusing to grant a top grade after a breach.
The approach also explains why apparently strong effort does not guarantee a strong finish. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four others, though less strongly.
There is also an important fairness qualification: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Publishing that difference helps readers interpret the ranking without pretending every surrounding condition was identical.

business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A benchmark should expose uncomfortable distinctions
Firmulate’s live company makes those distinctions observable. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn agent performance into an ongoing management record rather than a staged chat exchange.
The wider lesson is not that partial credit makes a benchmark soft. Here, it makes the test more honest. The score separates awareness from investigation, investigation from execution and execution from trustworthy execution. It also makes polished but unfinished work visible instead of forcing it into the same category as total inaction.
That evidence base extends to 242 real, unedited management decisions used in Firmulate’s model-guessing quiz. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For technology buyers, the important question is not whether an AI can produce a convincing answer. It is whether the system reads what matters, finishes what it starts and remains trustworthy when authority is ambiguous. A score of 26 for doing nothing is not a loophole. It is the control that makes every higher result easier to understand—and every near-perfect result worth examining rather than simply applauding.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI trust and reliability testing
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
