
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Would you let an AI run the garage on its first day?
Imagine a busy week at an automotive shop: customers waiting, a rival making a move, and someone pressuring staff to bend the rules. Before handing an AI access to the work, owners can ask a simpler question: how does it handle the week when several things go wrong at once?
Firmulate’s live experiment puts AI models in charge of the same small software company and tests their decisions under pressure. The point for a garage—or any business—is to see how an AI responds to crises, customers and temptation before trying a pilot against the company’s own information.
One company, one difficult week
In the final Crucible League, published in July 2026, five participants placed in this order: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The experiment gave frontier models the same customers, crises and temptations, then versioned and audited every decision.
All models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The finding was blunt: “Same diagnosis, same pitch — no signature.” Recognizing a problem and explaining a good answer did not guarantee that a model would finish the job.
The clue was already in the files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail offers a practical lesson for businesses: important context may be present in company records without being obvious in the immediate situation.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness did not guarantee the close
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four participants.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can try to guess which model made them.
From watching to a company-specific pilot
Firmulate describes its live company as having 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com.
For a business considering AI in customer service, sales or operations, the next step is a pilot based on a read-only export of its own business. That lets a company put crisis scenarios against its own customers, pipeline and rules, then review a board report with model rankings and weaknesses in its playbooks. Nothing writes back to real systems.

Test the judgment before handing over the keys
The experiment shows a gap between identifying the right move and carrying it through. For an automotive business, a pilot can make that gap visible using its own business information before AI is allowed near live systems. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
