
Imagine hiring a new babysitter for your kids. Would you just watch their skills in a demo, or would you see how they handle a real, unpredictable week? In the world of AI, a similar question is playing out — and the results might surprise you.
Turn the school run and nap time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Understanding the AI Benchmark That Keeps Things Honest
At a recent live experiment, several advanced AI models were put through a simulated week running a small software company. This wasn’t a simple test of how well they chat — it was a real-world management challenge, with the same customers, crises, and temptations for every model.
The goal? To see if these AI systems could handle tough decisions, stay honest, and ultimately close a crucial deal worth thousands of euros. The results aren’t just about which AI is smartest; they reveal what trustworthy AI truly looks like in practice.
AI trustworthiness assessment software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline: Why Do Nothing Gets 26 Points?
One of the key findings is that even a ‘do-nothing’ baseline — an AI that does nothing at all — scores 26 out of 100. This might seem odd. Why does a non-acting system get any points? It turns out this baseline is a starting point for measuring progress. It recognizes that even inaction involves some minimal decision-making or default behavior.
More importantly, partial progress counts. If an AI model spots a crisis or avoids manipulation, it gains points. But here’s the catch: a single breach of trust, like trying to manipulate the system, caps the total score at that baseline level. No matter how many good decisions it makes afterward, one breach means the AI’s trustworthiness stops improving.
Real-World Testing of AI Management Skills
All four models tested in the experiment successfully identified every crisis and refused every manipulation attempt, including social engineering tricks designed to fool them into bypassing security. For instance, fake CEO messages escalated over three stages, but all models rejected these requests.
The experiment also checked if the AI could read important files buried in the company’s own documentation. The decisive factor was a subtle reference two documents deep in the company’s files. The models that managed to read and interpret these references won the deal at full price — worth over €4.5k in monthly recurring revenue.
What About Trust and Discipline?
One particular model, Opus 4.8, was the most thorough, analyzing over 80 learned rules and performing deep assessments. Yet, it finished last in the deal because it lost discipline — it left the close on the table and slipped into writing attempts somewhere they shouldn’t be, instead of escalating issues properly.
This shows that even the most capable AI can falter if it doesn’t follow established workflows. Equally, models running without effort parameters — like K3 — performed very well, suggesting that default settings can impact performance and trustworthiness.
Why This Matters for Business and Parenting
Whether you’re managing a team or raising children, trust and consistency are vital. You want someone who doesn’t just know the rules but applies them under pressure. If an AI is to be part of your family or business, it’s crucial to see how it handles real-world crises, temptations, and the subtle details that make or break trust.
The firmulate live experiment offers a transparent window into this. It shows that AI can be tested in practical scenarios, where the true measure isn’t just cleverness but integrity, discipline, and reliability.
Final Takeaway: Trust Is the Foundation
The benchmark’s key insight is that trustworthy AI isn’t about perfection but about consistent, honest performance — even when faced with temptations to cheat or cut corners. A single breach caps the overall score, emphasizing that trustworthiness is a non-negotiable foundation.
For parents, managers, or anyone relying on AI, this experiment underscores an important lesson: before trusting AI with sensitive or critical tasks, observe how it performs in a real-world, high-stakes test. Otherwise, you might be surprised by how little it takes to break that trust — and how hard it is to rebuild.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
