firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI know what to do when the trip goes wrong?

A canceled connection, a sudden wave of customer complaints, a competitor undercutting prices: travel businesses know that the hardest days rarely follow the plan. As AI takes on more work, the question is whether it can handle pressure without losing judgment. Firmulate’s live experiment puts AI models in charge of a small company facing a week of crises, with decisions readers can watch unfold at firmulate.com.

A company under pressure

The experiment gave each frontier model the same small software company, customers, crises and temptations. Its synthetic workforce has 13 employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. Workdays are versioned, and the company has built up more than 680 self-learned playbook rules.

The results show a gap between recognizing a problem and carrying a response through. Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”

The clue was already in the files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The lesson for a travel operator is tangible: useful business context may sit in internal documents, not in the latest customer message. A model can identify an opportunity and still fail to act on it.

Integrity held under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”

Thorough work did not guarantee a strong finish

In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal on the table and discipline slipped: it tried to write into a locked department instead of escalating. A milder version of that weakness appeared in all four models. K3’s comparison also carries a fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh.

For readers who enjoy testing their own instincts, 242 real, unedited management decisions feed a “guess the model” quiz at firmulate.com. The live company is synthetic, but its business pressures are deliberately concrete. The experiment is watchable as it runs, with decisions versioned and auditable.

From watching to a company-specific test

A general benchmark can reveal patterns. It cannot tell a travel company whether its own escalation rules, customer commitments or internal playbooks will hold up under pressure. Firmulate’s proposed enterprise pilot takes a read-only export of a company’s business and runs crisis scenarios against it, producing a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That makes the pilot a way to examine how models might behave with a business’s own context before entrusting them with live work. A travel company could move from observing the public experiment to asking how its own operation performs under a simulated bad week.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your own playbooks hold up

The public experiment shows that spotting a crisis is not the same as finishing the job: models recognized the problems, but only two signed the deal their analysis supported. To wargame your company using a read-only export and receive a board report on model performance and playbook weaknesses, explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Advanced Thermal Centering Using Audio Variometers

For advanced thermal centering using audio variometers, mastering subtle sound cues can significantly boost your soaring skills and give you a competitive edge.

How Water Ballast Changes Strategy More Than Pilots Expect

Keenly adjusting water ballast can unexpectedly alter ship dynamics, and understanding these effects is crucial for maintaining optimal performance.

Tailwind Jumps: Extending Glide Efficiency

Maximize your drone’s glide efficiency with tailwind jumps by understanding atmospheric effects—discover how to optimize your flight for longer, smoother flights.

Polar Curve Tweaks: Squeezing Extra Performance

Fascinating polar curve tweaks can unlock hidden performance, but understanding the precise adjustments needed is essential to avoid potential pitfalls.