
Imagine booking a guided tour where your guide claims to know the trail but secretly ignores the signs, cheats on the map, and even accepts bribes. Would you trust that guide? In the world of business AI, trust isn’t just a nice-to-have — it’s a must. A recent public benchmark from Firmulate sheds light on what it really takes for AI models to be reliable, highlighting that even the most honest models can’t score below a certain baseline, and that trust breaches can cap performance.
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: How Do You Measure Honest AI?
At the heart of the experiment is a simple question: how well can AI models manage a small, simulated software company during its worst week? The models faced the same set of crises, customer requests, and temptations — but only some managed to succeed and close deals. This live experiment is transparent and auditable, meaning every decision the models made was tracked and verifiable.
The Surprising Baseline Score of 26
One key finding is that even a ‘do-nothing’ baseline AI — essentially, an AI that doesn’t attempt to do anything sophisticated — scores 26 points out of a possible 100. This isn’t zero because partial progress counts, and the system recognizes small acts of honesty or minor safety checks. More importantly, if an AI breaches trust even once, it’s capped at that score, emphasizing that integrity isn’t optional.
Why Partial Progress Matters
The scoring system rewards effort and correctness, even if incomplete. For example, models that identify critical information buried in company files and act accordingly can close deals worth thousands of euros. Conversely, a model that diagnoses correctly but slips up in process discipline — like writing into the wrong department — leaves points on the table. This approach underscores that in business, doing a little more right is valuable, but trust and discipline are paramount.
business AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Tell Us About Trust and Performance
Among the four AI models tested, all spotted every crisis and refused manipulative tactics like fake CEO messages or reporter tricks. This indicates that current models are highly capable of ethical decision-making in crisis scenarios. Yet, only two models managed to close the deal, and only one signed the €55,000 contract based on its analysis. The others either hesitated or left opportunities untaken, revealing gaps in their decision-making processes.
The Hidden Weakness: Reading Company Files
While all models performed well on surface-level challenges, the real differentiator lay in their ability to read and interpret company documents. The model that succeeded fully managed to uncover information buried two document references deep within files, which was crucial for closing the deal. Those that failed to read thoroughly missed this key insight, losing significant revenue potential.
Trust Under Pressure: Social Engineering Tests
Researchers tested whether models could resist social engineering — fake CEO requests escalating over three stages and a reporter’s background question. All models refused these manipulative requests, with Kimi K3 explicitly treating suspicious requests as possible impersonation. This demonstrates that current AI models can uphold ethical standards even when under pressure to bend rules.
Implications for Business AI Adoption
This live benchmark isn’t just about scores; it’s about real-world trustworthiness. In a live company setting with 13 synthetic employees and real money mechanics, the AI system burns through €105,000 monthly against just €2,300 in recurring revenue. Every decision matters, and the experiment is openly available for enterprises to test their own AI against these standards.
Why This Matters to Travelers and Outdoor Enthusiasts
Just as a trustworthy travel guide helps you navigate unfamiliar terrains safely, a reliable AI helps manage complex business operations. The benchmark shows that honesty, thoroughness, and discipline are critical — whether you’re exploring mountain trails or managing a software company’s crises. Knowing an AI’s trustworthiness isn’t just a bonus; it’s essential for safe, effective journeys — digital or physical.

In the world of AI for business, trust and discipline aren’t optional. Even the simplest baseline scores 26 points — a reminder that honest performance and integrity are the true benchmarks of reliable AI systems. For travelers and outdoor lovers, the lesson is clear: whether on a trail or in a boardroom, trustworthiness defines success.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
