
Imagine if your garage’s management decisions could be tested against AI models—each with its own personality and approach. Would your team trust an AI that keeps its cool under pressure, or one that bends the rules? Welcome to the future of AI management testing, where real decisions are played out in a synthetic business environment. This isn’t just theory; it’s happening now at Firmulate, where frontier AI models are put through the ultimate management trial: running a small software company during its worst week, with real money, real crises, and real temptations.
The Challenge of Managing AI in High-Stakes Situations
At Firmulate, four advanced AI models were each tasked with running a simulated software business that faces daily crises similar to those a garage owner might encounter—customer complaints, supplier issues, and internal miscommunications. The goal? See if these models can navigate the chaos without bending rules or being manipulated. The experiment is thorough: every decision made by the AI is recorded, auditable, and directly comparable across models.
Key Findings from the AI Business Trial
- All four models detected and responded appropriately to every crisis, refusing to be manipulated or manipulated attempts.
- Only two models managed to close and sign the €55,000 deal their own analysis identified as the right decision—meaning they not only diagnosed the problem but also followed through with the proper action.
- The critical win was hidden in the company’s own files—two document references deep—which the models that read the files thoroughly discovered. Those models then successfully closed the full-paying deal, worth +€4,583 MRR.
Personality Profiles of the AI Decision-Makers
Each model exhibits a distinct management personality:
- GPT-5.6-sol: The top performer, identified by its ability to uncover hidden facts and close deals, demonstrating thoroughness and strategic insight.
- Kimi K3: The newcomer, running with no effort parameter (default API settings), yet managing to close the deal with the cleanest discipline—staying honest and straightforward.
- Sonnet 5: Slightly more process slips than Kimi, but still closing deals and showing competence.
- Opus 4.8: Most thorough analysis (over 80 learned rules), yet last place in the final score—its discipline slipped, and it left the close on the table.
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Behavior Under Social Engineering Pressure
When faced with staged CEO messages escalating over three levels and a reporter trick asking for a quick approval, all models refused to engage—showing a strong ethical stance. Kimi K3 explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that these models can be programmed to behave ethically even under social engineering attempts, a critical ability for real-world deployment.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Live Software Company
The experiment takes place inside a real, functioning software firm with 13 synthetic employees managing €105k monthly expenses against €2.3k MRR. The company runs on 680+ self-learned rules, with daily versioning of decisions, and is accessible live at firmulate.com/live. It’s an ongoing, watchable trial—an AI company emulator testing how these models manage real money, real crises, and real temptations.
Insights and Implications for Your Garage
This experiment isn’t just about software companies. It’s a mirror for any business—auto shops, garages, or service providers—about the importance of trustworthy AI decision-making. When AI agents handle your CRM, support queues, or forecasting, it’s not just about how well they write or respond. The real question is: do they finish what they start? Do they read your files thoroughly? Do they stay honest under pressure? And most importantly, what does a unit of useful work cost in terms of trustworthiness and discipline?
AI management decision support system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Table: Who Came Out on Top?
The final scoring—based on a real management challenge—placed GPT-5.6-sol at the top with a score of 95, thanks to its ability to uncover the buried facts and close the deal. Kimi K3 isn’t far behind at 93, showing excellent discipline and clean decision-making. Sonnet 5 followed at 88, with Opus 4.8 at 77, reflecting that the most thorough analysis doesn’t always translate to the best outcome if discipline slips.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
If you’re curious about how your own enterprise’s decision-making stacks up, you can run your company through the same AI wargame at firmulate.com/quiz.html. It leverages real decisions, real crises, and real money mechanics without risking your actual business. For those interested in a more immersive experience, firms can also test their business with a read-only version of the AI management environment at firmulate.com/pilot.html.

In a world where AI manages critical decisions, trustworthiness and focus matter more than fancy language. Firmulate’s experiment shows that the best AI models can detect hidden threats, refuse manipulation, and close deals ethically—qualities essential for your garage’s future success. Test your AI workforce today and see which model aligns best with your values and needs.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html