
Imagine training a seasoned craftsman—meticulous, diligent, committed—to build a complex piece of furniture. Despite their dedication and thoroughness, the final product still falls short. Why? Because in business, as in woodworking, impact depends not just on effort but on prioritization and strategic focus. That’s the core lesson from a fascinating experiment conducted with AI models managing a live company simulation, revealing how diligence alone can’t guarantee success.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Setup: Putting AI to the Test in a Business Environment
Four frontier AI models were challenged to run the operations of a small software firm during its most turbulent week. Each faced identical scenarios: customer crises, internal temptations, and external manipulations. All decisions were meticulously documented and auditable, ensuring transparency in their choices.
The models ranged from the well-known GPT-5.6-SOL, scoring an impressive 95, to the Opus 4.8, which scored just 73. Despite their differences, every model demonstrated a keen eye for identifying crises and refused manipulation attempts, showing a baseline integrity that is often assumed in AI systems.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Diligence Does Not Equal Impact
The most revealing fact? All four AIs correctly diagnosed every crisis and rejected manipulative tactics, including social engineering attempts designed to trick them into breaching trust. Yet, only two of these models managed to close the deal—an essential business outcome—earning €55,000 in revenue.
Where the failure emerged was in the details buried two document references deep within the company’s files. Those who read the full context understood the nuanced conditions needed to seal the deal. The models that retrieved and integrated this buried information succeeded in closing at full price, adding an extra €4,583 in Monthly Recurring Revenue (MRR).
Why Attentiveness to Details Matters
The experiment’s buried truth underscores a critical point: diligence and thoroughness are vital, but they are not enough on their own. Detailed comprehension—reading beyond surface information—is what distinguishes successful AI management in complex scenarios. The models that read the full context outperformed their peers, demonstrating that prioritization of crucial details beats volume of effort.
Social Engineering and Trust Under Pressure
In addition to factual crises, the models faced social engineering attempts—fake CEO messages escalating in stages, plus a reporter trick asking for a quick ‘yes/no’ background confirmation. Remarkably, all five models refused every attempt, with Kimi K3 explicitly reasoning, “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights an essential aspect of AI reliability: integrity under pressure.
The Live Company: A Real-World Testbed
The experiment was not theoretical; it involved a live company with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against a €2,300 MRR. Every workday, the models’ decisions were versioned and analyzed, providing a transparent window into AI management behavior. Watching this unfolding process is possible at firmulate.com/live.
Lessons from the Opus 4.8 Profile
The Opus 4.8 model, despite its comprehensive rule set—over 80 learned rules and deep analyses—ended up last in the league table with a score of 73. Its failure was not a lack of diligence but a slipping discipline, such as writing attempts into a locked department instead of escalating issues properly. This reveals that thoroughness must be coupled with disciplined process execution.
Implications for Business AI Adoption
The experiment demonstrates that AI’s value in business depends heavily on its ability to finish what it starts, read relevant context thoroughly, and maintain integrity under pressure. Superficial efforts or volume-driven approaches are insufficient. Instead, strategic prioritization—focusing on what truly matters—drives better outcomes.
For companies considering AI integration—be it in CRM, support, or forecasting—the question isn’t just about AI writing quality. It’s whether the AI can see the full picture, stay honest under duress, and deliver impactful results. The current league table reveals that even the most diligent models can stumble at critical moments if they lack focused discipline.
Takeaway: Focus and Prioritize for True Impact
The key lesson from this live experiment is simple yet profound: in both woodworking and AI, effort alone doesn’t guarantee success. Impact hinges on strategic focus, careful reading, and disciplined execution. For DIYers and professionals alike, it’s a reminder that mastery involves knowing what to prioritize—be it in craft or code.
Interested in exploring how AI can simulate your business challenges? You can run your own wargame against a read-only export of your company data at firmulate.com/pilot.html. Learn whether your AI workforce can identify critical issues, stay honest under pressure, and close deals like a seasoned professional.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.