
Imagine a small business that operates with no human employees, yet faces the same crises and tough decisions as any company — and it’s all happening live online. For parents and families juggling busy days, it’s a reminder that even in the world of high-tech, survival often depends on integrity, discipline, and smart decision-making. Welcome to the daring experiment by Firmulate, where artificial intelligence models run a small software company through its worst week — and you can watch every move unfold.
The Live Experiment: An AI Company in Action
At the core of this groundbreaking project is a completely virtual company, managed by 13 synthetic employees powered by cutting-edge AI models. Every decision, crisis, and temptation is simulated with real financial stakes: the company burns through €105,000 each month but earns only €2,300 in monthly recurring revenue. It’s a high-wire act of resource management, honesty, and strategy, all played out in a publicly accessible online platform at firmulate.com/live.html.
What Does the AI Company Do?
The company runs a series of daily operations typical of small software firms — handling customer crises, negotiating deals, and making management decisions. Every day, the AI models are faced with the same challenges, from customer support issues to ethical temptations like manipulating data or bending rules. These models are tested against a set of 680+ self-learned rules and are able to analyze internal documents for critical insights, often winning deals at full price when they read the company’s hidden files.
How Do the Models Perform?
Four different frontier AI models participated in the experiment, each with varying capabilities and discipline. They all spotted every crisis and refused every attempt at manipulation, including social engineering tricks like fake CEO messages or reporter tricks. Yet, despite their honesty and awareness, only two of them managed to close the lucrative €55,000 deal that their own analysis identified as the best course of action. The other two either hesitated or left the deal on the table, revealing how discipline and thoroughness can influence outcomes.
What Are the Hidden Weaknesses?
The real vulnerability did not lie in customer interactions but in internal documentation. The models that read the company’s files uncovered critical details that led to closing the deal at full price — a difference worth over €4,500 in monthly recurring revenue. This shows that access to comprehensive data is crucial, and models that ignore internal knowledge might miss key opportunities.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond AI: The Human-Like Dilemmas
Interestingly, the models were tested against social engineering scenarios. Fake messages from a supposed CEO escalated through three stages, and attempts to persuade the AI to bypass approval processes were met with unwavering refusal. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This high level of resistance demonstrates AI’s potential to uphold integrity even under pressure, a vital trait for future automated decision-makers.
The Company’s Real Challenges
Despite the robust decision-making, the virtual company is not turning a profit. It is losing €105,000 each month against a tiny revenue of €2,300 — a stark illustration of how difficult it is to run a business without human oversight, especially when operating in a transparent, high-pressure environment. Every move is versioned and logged, and the system is designed for transparency, allowing anyone to see how decisions are made and whether rules are followed.
The Performance League
The models are ranked based on their scores, with gpt-5.6-sol leading at 95 points, having uncovered internal details and closed the deal. Kimi K3 follows closely at 93, demonstrating excellent discipline and decision quality. The other models scored 88 and 77, respectively, often leaving opportunities or making process slips that prevented deal closure. These rankings highlight how different AI models handle the same complex, human-like challenges.
Why This Matters for Families and Businesses
The firmulate experiment underscores a vital point: As AI begins to assist or replace parts of your family’s daily life — from managing schedules to customer service — the question isn’t just about how well it writes or responds. It’s whether it can see the full picture, make honest decisions, and stay disciplined under pressure. In a world where AI may touch your CRM, support queue, or forecast tools, understanding its ability to finish what it starts and act ethically is crucial.
Take Action and See for Yourself
You can watch this entire company’s daily struggles unfold online, see the decisions it makes, and even try your hand at guessing which AI model made each move at firmulate.com/quiz.html. For enterprises eager to test their own AI systems, there’s also a read-only export feature that lets them run the same scenarios without risking real data or systems — visit firmulate.com/pilot.html to learn more.

This real-time online experiment shows how AI can simulate running a company, revealing strengths like honesty and decision quality, but also exposing weaknesses that matter for all businesses. For families and anyone interested in the future of automation, it’s a live glimpse into how AI might handle real-world challenges — with transparency and accountability on full display.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html