
Imagine a business where every decision is made by AI, yet it operates in the red, battling daily crises while the world watches. This is not fiction, but the reality of a pioneering live experiment in building an AI-powered company in full public view. For homeowners and smart home enthusiasts, this story unveils how AI’s capabilities—and limitations—are tested in the harshest conditions, shaping the future of automated management and decision-making.
The Live Experiment: A Company Without Employees, Yet Full of Life
At the core of this groundbreaking project is a small, synthetic company powered entirely by AI models, operating with 13 simulated employees. Every weekday, the company faces real crises—customer issues, financial decisions, and strategic dilemmas—that are logged, analyzed, and responded to by AI. The entire process is publicly accessible at firmulate.com/live, allowing anyone to watch the company’s daily struggles unfold in real time.
AI-powered smart home security system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Complex Challenges in a Controlled Environment
The experiment pits four different frontier AI models against the same weekly challenges. Each is tasked with running the business through its worst week, with identical crises, customer requests, and temptations to manipulate or cheat. Every decision the models make is versioned and auditable, providing a transparent view into their reasoning processes.
smart home automation hub
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Performance and Surprising Outcomes
While all four AI models managed to identify every crisis and refused manipulation attempts consistently, only two of them managed to close a critical deal worth €55,000—a significant revenue target. Interestingly, the decisive factor wasn’t just their crisis management but a buried piece of information hidden deep in the company’s files, which the models that read and understood the documentation uncovered and used to secure the deal. This emphasizes that reading and comprehension are vital skills for AI in business contexts, beyond surface-level decision-making.
AI personal assistant device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Human-Like Deceptions and AI Integrity
The experiment also tested the models against social engineering tactics, such as fake CEO messages and reporter tricks. All five AI models refused to be manipulated, with Kimi K3 explicitly treating suspicious requests as potential impersonations, demonstrating a strong capacity for integrity under pressure.
home automation security camera
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Cost of AI Management
Despite their decision-making prowess, the live company is far from profitable. It burns €105,000 every month, earning only €2,300 in monthly recurring revenue. The public cash countdown illustrates the precarious financial situation, highlighting the immense challenge of building sustainable AI-driven enterprise models.
Lessons from the Front Lines
Deep analyses reveal that even the most disciplined AI models can slip under stress. For example, Opus 4.8, the most thorough participant with over 80 learned rules, failed to close the deal, leaving it unexecuted, with discipline slipping during the final stages. The same weakness appeared, albeit less severely, across other models. This indicates that thoroughness does not always translate into successful execution, especially under pressure.
Implications for Home and Smart Technologies
For homeowners and users of smart devices, this experiment offers a cautionary tale and a sign of what’s to come. If AI can manage a business, handle crises, and maintain integrity in a high-stakes environment, similar principles will soon apply to smart homes, appliances, and personal assistants. The question is not whether AI will write well or converse convincingly, but whether it can finish what it starts, stay honest, and adapt under real-world pressures.
Building Trust and Understanding AI Limits
The leaderboard shows that the highest-scoring model, gpt-5.6-sol, not only identified and closed the deal but also uncovered the crucial hidden document. Meanwhile, others performed almost as well but faltered at the final hurdle. This underscores that trustworthiness and thoroughness are key qualities for AI in business or home automation—especially when lives and significant money are involved.
Experience It Yourself
Interested in seeing AI decision-making in action? You can observe this ongoing experiment at firmulate.com/live. Additionally, the platform offers a quiz at firmulate.com/quiz to test your management intuition against real decisions made by AI models. For enterprises, there’s an option to run similar simulations against your own business, ensuring your AI tools are battle-tested before deployment in critical environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html