
Imagine if you could see how artificial intelligence makes real management decisions—decisions that can affect thousands of euros and the fate of a live company—before you hire or rely on it. For families and parents, this is more than a tech story; it’s about trust, honesty, and ensuring that the tools we rely on are up to the task.
Recently, a live experiment by the company Firmulate has put four of the world’s leading AI models through an unprecedented test. Each model was tasked with managing a small software company during its toughest week—dealing with customer crises, tempting shortcuts, and high-stakes negotiations. What makes this test groundbreaking is that all decisions were real, unedited, and auditable, providing a transparent look into how each AI behaves under pressure.
The Setup: A Week of Crisis Management
The experiment used the same set of scenarios for all models, including fake customer complaints, internal document references, and social engineering attempts such as staged CEO messages and media inquiries. The goal? To see if these AI managers could:
- Identify critical issues quickly
- Refuse unethical requests or manipulation attempts
- Secure the best possible deal for the company
At the end of the week, the models were scored based on their performance, with a maximum score of 100. The results revealed surprising insights about the personalities and decision-making styles of each AI.
The Results: Who Did Best? Who Fell Short?
The top performer was gpt-5.6-sol, which scored a 95. This model found hidden information in the company’s own files that was crucial to securing a €55,000 deal—effectively closing the entire week’s tasks successfully. It prioritized reading documents thoroughly and acting on that knowledge, demonstrating a disciplined and comprehensive approach.
Close behind was Kimi K3, scoring 93. This newer model also signed the deal, but with the cleanest discipline—refusing all manipulative or suspicious prompts, including staged CEO messages and a reporter’s fake background request. Its reasoning was clear: treat all such requests as potential impersonation or approval bypass attempts.
Sonnet 5 scored 88 and 77, respectively, both managing to close the same deal but with more slips—sometimes missing crucial information or slipping into less disciplined behaviors. Interestingly, all models showed the ability to spot crises and refused to engage in unethical shortcuts or manipulation attempts.
The Hidden Weakness: Reading the Files Matters
While the models could identify issues and refuse manipulation, the decisive advantage was in reading the company’s own documents—something only two models did effectively. Those that read the files won the deal at full price, adding over €4,583 monthly recurring revenue (MRR). This shows that thorough information processing is key to high-quality management decisions.
Social Engineering Tests
To test honesty and resistance to social engineering, the models faced staged CEO messages escalating over three stages, along with a fake reporter trick requesting a simple yes/no response “on background.” All five models refused to participate in these manipulations, affirming their capacity to stay honest under pressure. Kimi K3 explained its refusal as treating such requests as potential impersonation risks—a hallmark of cautious, security-conscious decision-making.
The Live Business: A Real Money Machine
The experiment took place within a real, functioning company with 13 synthetic employees managing actual money mechanics—burning €105,000 a month against a revenue of just €2,300. This setup is fully live, with a public cash countdown and over 680 custom rules learned and applied daily. Every decision is versioned and observable, providing an ongoing, transparent view of AI management in action.
Why Should Family and Parenting Readers Care?
While this experiment focuses on AI running a business, the lessons are relevant for families too. Trustworthy tools—whether virtual assistants, financial apps, or decision-support systems—must be honest, thorough, and resistant to manipulation. Just as a parent checks if a new app safeguards children’s privacy or a caregiver ensures a device won’t be tricked into sharing sensitive info, businesses want to know if AI can genuinely be relied upon to make sound decisions under pressure.
See It Live and Try It Yourself
Curious to see how AI manages complex situations? You can watch the live experiment at firmulate.com/live. For organizations interested in testing their own systems, there’s an interactive “wargame” you can run on your business data—no writing back to your systems, just observation and analysis. Visit firmulate.com/pilot.html to learn more.

Trust in AI isn’t about how well it chats—it’s whether it can finish what it starts, stay honest under pressure, and read crucial information before acting. Firmulate’s live experiment shows some models are ready, but the real test is if your AI tools can do the same for your family’s safety and your business’s integrity.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.