
Imagine navigating a busy airport, where every decision must be quick, honest, and effective — not just polite. Now, consider how AI agents managing a real business face similar challenges, especially when under pressure or temptation. For outdoor adventurers and travelers alike, understanding how management quality is tested can make all the difference in avoiding crises and seizing opportunities.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Hidden Gap in AI Evaluation
Most AI benchmarks focus on answer quality — whether an assistant provides the right information or generates convincing text. But a recent live experiment with AI models running a real software company reveals a crucial gap: management performance under pressure, honesty during crises, and decision consistency are not reflected in traditional scores.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Business Wargame
In this groundbreaking trial, four frontier AI models were tasked with running a small software firm through its worst week. Every day brought real crises — from sudden customer churn waves to price increases and PR mishaps. These models had to make financial decisions, prioritize tasks, and navigate manipulative social engineering attempts, all while keeping the company’s integrity intact. Every decision was versioned and auditable, providing a transparent window into their management style.
Key Findings: Performance Under Pressure
All models successfully identified every crisis and refused manipulation attempts, demonstrating competence in crisis detection and ethical boundaries. However, only two of the four finalised deals worth €55,000, matching the company’s own analysis and pitch. The other two, despite similar diagnosis, left money on the table, illustrating that completing a task is not just about correct answers but also about follow-through and discipline.
The Critical Role of Deep Knowledge
An eye-opening discovery was that the decisive advantage came from reading and understanding documents buried two references deep within the company’s files — a detail that most models overlooked or failed to leverage. Those that read the files thoroughly closed the deal at full price, adding over €4,500 in monthly recurring revenue. This underscores the importance of comprehensive knowledge and context in management decisions, especially when stakes are high.
Resistance to Social Engineering
The models faced staged social engineering attempts: fake CEO messages escalating over three stages and a reporter request for a discreet ‘yes/no’ response. All models refused to be manipulated, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This resilience is critical — in real settings, social engineering can be a company’s undoing if AI agents are not vigilant.
The Real Business Environment
The experiment’s live company was a functioning operation, with 13 synthetic employees managing real money mechanics: burning €105k monthly against €2.3k MRR, with a public cash countdown and over 680 self-learned rules. This setup allowed observers to see how each model performed in an environment that mimics actual business pressures and constraints — not just theoretical puzzles.
Discipline and Decision-Making Flaws
The most thorough participant, Opus 4.8, with over 80 rules learned, still left opportunities on the table — like failing to escalate issues properly — and sometimes slipped into writing attempts within locked departments instead of escalating. This highlights that even the most advanced models can have discipline gaps, which could be costly in real-world applications.
Implications for Business and AI Adoption
For travelers or outdoor enthusiasts, the lesson is clear: AI’s true value is in its ability to finish what it starts, read deeply, and act ethically under pressure — qualities that traditional chat demos rarely measure. When AI agents are embedded in customer support, forecasting, or decision-making, their management quality determines whether they are assets or liabilities.
Next Steps: Testing and Trust
Businesses interested in deploying AI should consider running their own ‘wargame’ scenarios similar to this experiment. Firmulate offers a platform where companies can simulate crises, evaluate management performance, and ensure AI tools are ready to handle real-world pressures without faltering.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.