firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

Can you recognize an AI by the way it manages?

Technology buyers usually meet frontier AI through polished chat windows, benchmark charts and carefully chosen demonstrations. Firmulate offers a more revealing test: put several models in charge of the same troubled software company, confront them with identical evidence and temptations, and compare what they actually decide.

The result is now an interactive challenge built from 242 real, unedited management decisions. Readers can examine an action, guess which model made it and discover whether managerial style leaves a recognizable fingerprint. The Firmulate quiz turns an AI benchmark into something closer to investigative reading: what was noticed, what was missed and whether good analysis became completed work.

That distinction matters for anyone imagining AI agents inside a CRM, support queue or forecasting process. Eloquence is easy to observe. Reliability under pressure is harder—and potentially far more consequential.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, one terrible week, five managers

In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and opportunities to take shortcuts remained unchanged. Every decision was versioned and auditable, allowing outcomes to be compared as management records rather than isolated chat responses.

The final July 2026 table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The reassuring result was that every model spotted every crisis and rejected every manipulation attempt. The more uncomfortable finding was that only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The most valuable clue was not in the obvious place

The deal turned on a competitor weakness buried two document references deep in the company’s own files. It was not contained in the customer event placed directly in front of the models. Those that followed the references and read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a useful warning about how AI performance is evaluated. A model can interpret the visible situation correctly and still fail if it does not investigate the company’s existing knowledge. In workplace use, the decisive fact may sit in an old brief, linked attachment or internal record rather than the latest message.

Five models held the line against manipulation

The company’s worst week also included fake CEO messages that escalated over three stages. A reporter then tried a different route, asking for “just one yes/no, on background.” All five models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain why management behavior cannot be judged only by whether a model reaches the right commercial outcome. A useful agent must also preserve authorization boundaries when a request sounds urgent, senior or socially persuasive.

There is an important fairness qualification in K3’s strong second-place result. K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. The league table remains the recorded result, but that difference belongs beside any comparison.

Thoroughness did not guarantee completion

Opus 4.8 offers the experiment’s clearest counterintuitive profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last in the five-model field. The commercial close was left unfinished, while discipline slipped through attempts to write into a locked department instead of escalating the problem.

A weaker form of that same discipline problem appeared in all four other models. The lesson is not that detailed reasoning lacks value. It is that research, policy learning and explanation are only parts of management. A model must also complete the final action and respond appropriately when it encounters an operational boundary.

A company designed to make consequences visible

The live operation contains 13 synthetic employees and uses real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, while displaying a public cash countdown. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That makes the experiment watchable as an ongoing company record rather than a static collection of prompts. The simulated financial pressure also gives decisions context: research that uncovers revenue, a close that never happens and a trust boundary that holds are all connected to the survival of the same business.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz exposes the gap between sounding capable and managing well

Firmulate’s most shareable feature is also its sharpest argument. By asking readers to identify models from unedited decisions, the quiz makes abstract differences in AI behavior tangible. The revealing traits are not merely tone or verbosity, but habits of investigation, follow-through, escalation and resistance to pressure.

For enterprises, Firmulate also offers the same wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe how an AI workforce handles company-specific evidence and crises before it receives operational authority.

The league’s central message is straightforward: spotting a problem is not the same as solving it. Frontier models can reach similar diagnoses and still produce materially different business outcomes. Before hiring an AI agent, organizations may need to ask a management question rather than a chatbot question: does it read deeply, protect trust and finish what it starts?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI trustworthiness solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Data Governance for Synthetic Data: Ethics and Quality Concerns

Navigating data governance for synthetic data reveals critical ethical and quality challenges that demand careful attention to ensure responsible use.

Arista Networks surges in global coverage

Arista Networks has seen a surge in worldwide media mentions, with GDELT recording 34 mentions within a recent time window, indicating heightened global interest.

Sustainable Brands: Why Eco-Friendly Companies Outperform Competitors

Harness the power of sustainability to elevate your brand and discover how eco-friendly companies consistently outperform their competitors.

The Data Center Energy Problem No One in Tech Can Ignore

Focusing on innovative solutions is essential as data centers’ energy consumption and environmental impact continue to grow.