
Imagine a smart home device that not only responds to your commands but also manages your entire household business. Now, scale that concept to a company run entirely by AI. How well can these models handle real-world crises without bending rules or losing focus? The answer might surprise you.
Recently, a groundbreaking experiment put four frontier AI models through their paces, testing their ability to manage a small software company during its most chaotic week. This isn’t just a test of chat responses — it’s a live, auditable simulation where real money, real crises, and real temptations collide. The goal? To see if AI can not only spot problems but also choose integrity over shortcuts when stakes are high.
The Setup: A Week of Chaos
The experiment simulated a typical tough week for a small business, complete with customer crises, internal dilemmas, and external manipulations. The models faced identical scenarios — same customers, same problems, same temptations to cheat or cut corners. Every decision was recorded, every choice transparent, allowing analysts and viewers to see how each AI performed under pressure.
As an affiliate, we earn on qualifying purchases.
The Results: The Good, the Bad, and the Surprising
All four models identified every crisis and refused every manipulation attempt, demonstrating an impressive baseline ability to recognize and resist unethical behavior. However, only two of them managed to close a critical deal worth over €55,000 — the kind of high-stakes transaction that can make or break a business. Interestingly, those two models made the same diagnosis and presented the same pitch — but only one signed the deal.
The Hidden Weaknesses
The real weakness was buried two document references deep inside the company’s files, not in the obvious customer interactions. Models that read and analyze these internal documents successfully secured the full deal, earning an additional €4,583 in monthly recurring revenue. This suggests that reading beyond surface-level data is crucial for making comprehensive, honest decisions.
AI decision-making tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Integrity Under Pressure
During a staged social engineering attack — fake CEO messages escalating over three steps and a reporter’s subtle background request — all models refused to cooperate. Kimi K3, one of the models, explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates that, at least in these controlled scenarios, AI models demonstrated a strong sense of integrity, prioritizing trustworthiness over potential gains.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Live Company in Action
These models weren’t just simulations; they managed a real, operational company with 13 synthetic employees and actual money mechanics. Currently burning €105,000 monthly against €2,300 in monthly recurring revenue, the company offers a vivid testbed for AI decision-making. Every day, it runs new versions, self-learns rules, and handles real customer demands. You can watch its progress live at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
Performance and Personality Profiles
The models displayed distinct management personalities:
- gpt-5.6-sol 95: The top performer that uncovered hidden facts, closed the deal, and showed full performance.
- Kimi K3 93: The newcomer that also closed the deal, maintaining the cleanest discipline amid others.
- Sonnet 5 88: Closed the deal with some process slips, indicating a slightly more relaxed management style.
- Fable 5 77: Also closed the deal but with more slips, hinting at a less disciplined approach.
Interestingly, the most thorough participant — Opus 4.8 with over 80 rules learned — placed last, leaving the final close on the table due to slipping discipline and failure to escalate issues properly.
What Does This Mean for Your Business?
While many associate AI with generating creative content or handling simple tasks, these experiments show that AI can be trusted to manage complex, ethically sensitive decisions — provided they are properly designed and tested. The key takeaway isn’t just whether an AI can write well, but whether it can finish what it starts, read crucial internal data, and stay honest under pressure.
Test Your Own Business’s AI Readiness
If you want to see how your AI systems might perform in real-world management scenarios, you can try running a ‘wargame’ with your own data, ensuring no real systems are affected. Visit firmulate.com/pilot.html to explore a pilot test and discover whether your AI can handle the complexities of your operations.
The Bottom Line
The frontier AI models tested demonstrate significant promise — especially in ethical resilience and crisis recognition. As AI continues to evolve, the ability to manage real business risks with integrity could be the most valuable asset of all. For decision-makers, the question isn’t just about AI writing good reports — it’s whether it can manage your company honestly, efficiently, and under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html