firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

What if your AI worker could read your files before making a decision?

Imagine having an AI that doesn’t just respond to prompts but truly understands the depth of your business documents—finding that crucial piece of information buried two references deep. This capability can make or break deals, especially in high-stakes environments where every detail counts.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Through Its Worst Week

Recently, a groundbreaking experiment tested four advanced AI models by placing them in a simulated environment mimicking a small software company’s most challenging week. Each AI faced the same customer crises, temptations to cheat, and manipulative scenarios. The goal was simple: see which AI could operate with integrity and find the hidden, decisive information buried deep inside company files.

Key Findings from the Competition

  • Performance Scores: The models scored from 73 to 95 points out of 100, with the top scorer being GPT-5.6-sol at 95 points and the lowest, Sonnet 5, at 88.
  • Crucial Discovery: All four models detected every crisis and refused manipulation attempts, demonstrating they could act ethically under pressure.
  • Winning the Deal: Only two models managed to sign the €55,000 deal, reflecting real-world business value. The others diagnosed the issues but failed to close, leaving money on the table.

The Hidden, Decisive Weakness

What separated the winners from the rest? The decisive weakness was not in answering questions but in reading and understanding complex internal documents. The models that read the company’s files thoroughly before responding were able to identify a crucial piece of information—hidden two references deep—that clinched the deal. This buried fact, worth an additional €4,583 MRR, proved that reading comprehensively can be the key to business success.

Social Engineering Tests and Ethical Boundaries

In a layered test of integrity, all models faced sophisticated social engineering—fake CEO messages escalating over three stages and a reporter trick asking for quick approvals. All refused, with Kimi K3 explicitly recognizing the risk of impersonation or bypassing approval processes. This demonstrates these models’ capacity to uphold trust and ethical standards, even under tempting pressure.

The Real Business: A Live Company Emulation

The experiment isn’t just theoretical. It’s being run on a live, synthetic company with 13 employees, real money mechanics, and rigorous self-learned rules. The company burns €105k monthly against €2.3k MRR, with every decision documented and auditable. Watch it unfold at firmulate.com/live.

Insights and Limitations

Among the models tested, Opus 4.8 was the most thorough, analyzing over 80 learned rules but still left a deal on the table due to discipline lapses—writing into a locked department instead of escalating. The most disciplined model, Kimi K3, with no effort parameter set, closed the deal successfully, showing that focus and integrity matter just as much as analysis depth.

What Business Leaders Should Take Away

For organizations considering deploying AI, the lesson is clear: it’s not just about how well the AI generates text or interacts in chat. It’s about whether it can read and understand your internal documents before acting, maintain honesty under pressure, and deliver tangible business outcomes. The ability to uncover buried facts that can seal deals or prevent losses is a measurable, decisive advantage.

The Future of AI in Business Decision-Making

As AI models mature, their capacity to read, comprehend, and trustworthiness will determine their true value. The current league table from the Crucible League highlights that models like GPT-5.6-sol and Kimi K3 are leading, with scores of 95 and 93 respectively, indicating they excel in these facets. Meanwhile, even the best can slip—Opus 4.8 scored 73, leaving room for improvement.

In practical terms, companies can simulate their own scenarios through tools like the pilot environment. Here, you can run your own business wargame, testing your AI’s ability to read and act before deployment, reducing risks and ensuring that your AI workforce performs ethically and effectively in real-world settings.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Key Takeaway

In AI-powered business decision-making, reading comprehension and integrity matter more than ever. The models that can sift through complex internal files and find crucial, buried facts are the ones that win deals, prevent losses, and build trust—turning AI from a chat tool into a strategic asset.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How Hard Is It To Learn To Sew – Ultimate Guide to Mastering the Craft

The journey to mastering sewing is filled with challenges and rewards; discover the secrets that will transform your skills and ignite your creativity.

How to Spot Fabric Damage Before It Gets Worse

Using careful inspection, you can identify early signs of fabric damage before it worsens, helping you preserve your clothes longer and avoid costly repairs.

Why Your Fabric Keeps Bleeding Color (and How to Stop Dye Transfer)

Keen to keep your fabrics vibrant and avoid color transfer? Discover essential tips to stop dye bleeding effectively.