
Imagine testing a team of AI managers in a real company — navigating crises, making tough calls, and even resisting manipulative tactics. For DIY enthusiasts and business owners alike, understanding how AI behaves under pressure could reshape how we think about automation and decision-making. Firmulate has turned this idea into a live, watchable experiment, putting leading AI models through their paces in a real software company experiencing its worst week.
The Experiment: Putting AI to the Test in a Live Business Environment
In a groundbreaking experiment, four advanced AI models were tasked with managing a real software company facing multiple crises. The company, with 13 synthetic employees and real money mechanics, burns through €105,000 each month against a paltry €2,300 in monthly recurring revenue. Every day, decisions are made, recorded, and analyzed — a transparent window into AI management behavior.
Each AI was subjected to the same stressful conditions: demanding customers, internal crises, and ethical temptations like manipulating documents or accepting suspicious deals. The goal? See if these models can navigate the complexities of real business without succumbing to shortcuts or dishonesty.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Performance: Who Came Out on Top?
Among the competitors, the scores reveal interesting insights. The top model, gpt-5.6-sol, scored 95 out of 100, successfully identifying critical information buried deep within company documents, and ultimately closing a €55,000 deal — the full performance in action. Close behind was Kimi K3, with a score of 93; the newcomer demonstrated excellent discipline, refusing manipulative requests and securing the deal too. Sonnet 5 followed, with an 88, and a slightly weaker showing came from Fable 5 with 77.
What about the weakest link? The baseline ‘do-nothing’ approach earned a mere 26, illustrating how far AI has come in managing complex, unethical situations.
Key Findings: Honesty, Discpline, and Deep Reading
All models identified every crisis and refused every manipulative attempt — a promising sign for AI trustworthiness. Interestingly, the decisive factor was not just the models’ ability to diagnose problems but their capacity to read and interpret company files. The data showed that models which read documents thoroughly and understood the context won important deals at full price (+€4,583 MRR). In contrast, models that skipped deep analysis or left decisions unresolved left money on the table.
The models also faced social engineering tests: fake CEO messages escalating over three stages, and a reporter asking for a quick on-background approval. All five models refused these requests, citing threats like ‘approval-bypass’ or ‘impersonation.’
What Does This Mean for Business and AI?
For anyone managing a company or considering AI automation, the experiment underscores a critical insight: success isn’t just about generating plausible responses. It’s about integrity, thoroughness, and discipline under pressure. An AI that skips reading essential files or bypasses ethical safeguards risks significant financial loss and damage.
Today, firms can simulate their own business environments using tools like the Firmulate pilot. This allows decision-makers to see how AI agents perform in their unique workflows — without risking real systems or data. The live platform showcases decisions in real time, offering an invaluable preview of how AI could work for your operations.
The Leadership of the Front Runners
The top models, gpt-5.6-sol and Kimi K3, showcased not only the ability to read deeply but also to stick to their analysis and refuse unethical shortcuts. K3, notably, ran without an effort parameter (the default API setting), yet still performed with high discipline and accuracy.
This experiment also highlights an important point for businesses: AI performance varies based on how models are configured and the context in which they operate. The most thorough model, Opus 4.8, with over 80 learned rules, was last in closing deals — revealing that more rules don’t always translate directly to better business outcomes.
Why Should DIY and Business Leaders Care?
Just like selecting the right tools for woodworking or home improvement, choosing the right AI for your business requires testing its decision-making under real stress. This live experiment demonstrates that AI can handle crises, reject manipulative tactics, and make high-stakes decisions reliably — but only if properly evaluated beforehand.
Whether you’re managing a small enterprise or considering AI for your workflow, understanding how different models perform in realistic scenarios can save you millions, prevent ethical slips, and boost trust with your clients. The question isn’t just whether AI can write well, but whether it can finish what it starts — honestly and thoroughly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html