
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Before you trust an AI with the family business, give it a hard week
Parents know that keeping a household running takes more than spotting a problem. Someone has to make the call, follow through and know when a request should raise a red flag. Businesses face the same test as they hand more work to AI. Firmulate’s experiment asks what happens when models have to manage a company through a crisis, not just talk about one.
A shared crisis, with real consequences inside the experiment
In the final Crucible League in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. A breach of trust capped a total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. The diagnosis and pitch were there; the signature was not. It is a gap that a polished chat answer can conceal.
The detail hidden in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. The experiment rewarded not only recognizing the situation but finding and acting on information the company already had.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s appeal for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of restraint matters when a confident message arrives at a busy moment.
Thorough work still has to reach the finish line
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail to keep in view when reading the ranking.
The public experiment continues beyond the league. The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. A separate quiz draws on 242 real, unedited management decisions, inviting readers to guess the model. The live experiment is watchable at Firmulate.
From watching to trying it against your business
For a business leader, the point is practical: an AI can identify the right problem and still fail to close the loop. A pilot lets an enterprise put its own playbooks and decision-making under pressure. Firmulate says the wargame can run against a read-only export of a company’s business, with nothing writing back to real systems. That gives teams a way to examine how models handle their customers, crises and internal rules before trusting them with live work.

Put your playbooks through a rehearsal
Firmulate’s results show why AI evaluation should include follow-through, judgment and trust under pressure—not just a convincing answer. Enterprises can run a pilot against a read-only export of their own business. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
