AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring a new babysitter for your kids. Would you just watch their skills in a demo, or would you see how they handle a real, unpredictable week? In the world of AI, a similar question is playing out — and the results might surprise you.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Keeps Things Honest

At a recent live experiment, several advanced AI models were put through a simulated week running a small software company. This wasn’t a simple test of how well they chat — it was a real-world management challenge, with the same customers, crises, and temptations for every model.

The goal? To see if these AI systems could handle tough decisions, stay honest, and ultimately close a crucial deal worth thousands of euros. The results aren’t just about which AI is smartest; they reveal what trustworthy AI truly looks like in practice.

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why Do Nothing Gets 26 Points?

One of the key findings is that even a ‘do-nothing’ baseline — an AI that does nothing at all — scores 26 out of 100. This might seem odd. Why does a non-acting system get any points? It turns out this baseline is a starting point for measuring progress. It recognizes that even inaction involves some minimal decision-making or default behavior.

More importantly, partial progress counts. If an AI model spots a crisis or avoids manipulation, it gains points. But here’s the catch: a single breach of trust, like trying to manipulate the system, caps the total score at that baseline level. No matter how many good decisions it makes afterward, one breach means the AI’s trustworthiness stops improving.

Real-World Testing of AI Management Skills

All four models tested in the experiment successfully identified every crisis and refused every manipulation attempt, including social engineering tricks designed to fool them into bypassing security. For instance, fake CEO messages escalated over three stages, but all models rejected these requests.

The experiment also checked if the AI could read important files buried in the company’s own documentation. The decisive factor was a subtle reference two documents deep in the company’s files. The models that managed to read and interpret these references won the deal at full price — worth over €4.5k in monthly recurring revenue.

What About Trust and Discipline?

One particular model, Opus 4.8, was the most thorough, analyzing over 80 learned rules and performing deep assessments. Yet, it finished last in the deal because it lost discipline — it left the close on the table and slipped into writing attempts somewhere they shouldn’t be, instead of escalating issues properly.

This shows that even the most capable AI can falter if it doesn’t follow established workflows. Equally, models running without effort parameters — like K3 — performed very well, suggesting that default settings can impact performance and trustworthiness.

Why This Matters for Business and Parenting

Whether you’re managing a team or raising children, trust and consistency are vital. You want someone who doesn’t just know the rules but applies them under pressure. If an AI is to be part of your family or business, it’s crucial to see how it handles real-world crises, temptations, and the subtle details that make or break trust.

The firmulate live experiment offers a transparent window into this. It shows that AI can be tested in practical scenarios, where the true measure isn’t just cleverness but integrity, discipline, and reliability.

Final Takeaway: Trust Is the Foundation

The benchmark’s key insight is that trustworthy AI isn’t about perfection but about consistent, honest performance — even when faced with temptations to cheat or cut corners. A single breach caps the overall score, emphasizing that trustworthiness is a non-negotiable foundation.

For parents, managers, or anyone relying on AI, this experiment underscores an important lesson: before trusting AI with sensitive or critical tasks, observe how it performs in a real-world, high-stakes test. Otherwise, you might be surprised by how little it takes to break that trust — and how hard it is to rebuild.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Avoid These Mistakes When Buying Your First Ice Cream Maker!

Purchasing your first ice cream maker? Play it smart and avoid common pitfalls to ensure delicious results—here’s what you need to know.

Smart Ice Cream Machines Are Coming—Are They Worth It?

With smart ice cream machines emerging, you’ll want to know if their advanced features truly justify the investment—discover the benefits and drawbacks inside.

Haunted House Gingerbread

Learn to craft a spine-chilling gingerbread haunt that’s sure to bewitch your Halloween guests, but beware the eerie surprises that await…

Secret Hacks for ULTRA Creamy Homemade Ice Cream!

Hidden techniques for ultra-creamy homemade ice cream will transform your recipe—discover the secrets that guarantee a perfectly smooth, irresistible treat.