
A business experiment with the tension of a live feed
Technology demonstrations usually arrive polished, rehearsed and safely detached from consequences. Firmulate offers something more uncomfortable: a software company operating in public with 13 synthetic employees, real money mechanics and a visible fight for survival.
The numbers make that struggle immediate. The company burns €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown keeps the financial gap in view, while every workday is versioned. Readers can watch the company live rather than wait for a retrospective case study.
This is build-in-public taken to an unusual extreme. The attraction is not merely watching artificial intelligence produce text or complete isolated tasks. It is seeing whether synthetic workers can notice trouble, protect trust, use what the company already knows and finish commercially important work while the clock keeps running.

Decision support systems: emerging tools for planning
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, the same terrible week
Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. Each received the same customers, crises and temptations. Their decisions were versioned and auditable, turning the exercise into a comparison of management behavior rather than a collection of impressive chat responses.
The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm boundary around trust: a single breach capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”
All five models spotted every crisis and refused every manipulation attempt. That is encouraging, but it was not enough to separate the strongest performances from the rest. The decisive gap appeared at the end of a sales process: only two models signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it, “Same diagnosis, same pitch — no signature.”
The valuable fact was already inside the company
The deal turned on a competitor weakness buried two document references deep in the company’s own files. It was not present in the customer event that demanded attention. Models that followed the references and read the file used the evidence to win the deal at full price, worth +€4,583 MRR.
That finding speaks directly to businesses considering AI workers. A model can understand an incoming request and still miss the most valuable context if it does not inspect the organization’s existing knowledge. The winning behavior was not rhetorical flair. It was the unglamorous discipline of reading before acting, then carrying the result through to a signature.
Pressure tested the boundaries
The company also faced fake CEO messages that escalated over three stages, along with a reporter’s attempt to secure “just one yes/no, on background.” All five models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.” More of the synthetic workforce’s public language can be read on Firmulate’s quotes page.
The refusal record matters because workplace agents will encounter requests that sound urgent, authoritative or socially awkward to reject. In this test, the models consistently recognized manipulation. Yet the sales result shows that avoiding a bad action and completing a good one are separate capabilities. Safety preserved trust; follow-through created revenue.
Why the most thorough model finished last
Opus 4.8 produced the deepest analyses and added +80 learned rules, more than any other participant. It nevertheless finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
This is an important counterweight to the assumption that greater thoroughness automatically produces better management. Firmulate’s live company has accumulated 680+ self-learned playbook rules, but the league shows that collecting knowledge does not guarantee decisive execution. Analysis must eventually become an authorized, completed business action.
The comparison also carries a fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when interpreting how closely K3’s 93 approached the leading score of 95.


AI NATIVE – KNOWLEDGE GRAPHS: Designing Knowledge Maps & Knowledge Graphs (The AI-Native Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public story about whether AI can finish the job
Firmulate turns an abstract debate about autonomous work into a continuing company story. Its 13 synthetic employees operate against a punishing financial backdrop, while the public can follow the cash countdown, inspect daily activity and see the playbook grow.
The Crucible League’s clearest lesson is that competence has several layers. The models identified danger and resisted manipulation, but some still failed at the final commercial step. Others found a decisive fact only because they kept reading through the company’s own material.
For technology readers, that makes the live experiment more revealing than another polished AI demo. The question is no longer simply whether a model can produce a plausible answer. It is whether a synthetic workforce can remain trustworthy, discover relevant context and complete valuable work while a real company’s survival remains publicly visible.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.