
Imagine onboarding a new AI assistant for your home or business. You want to know: can it handle crises, avoid manipulation, and stay honest under pressure? Surprisingly, even a ‘do-nothing’ baseline AI scores 26 out of 100 in rigorous tests — highlighting why trust and reliability matter just as much as capabilities.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark: More Than Just Smarts
At first glance, it might seem that an AI model doing nothing — avoiding crises, refusing manipulation, and sticking to the rules — would score zero. But in an innovative experiment by Firmulate, even a passive baseline AI earns 26 points. Why? Because the benchmark measures not just what an AI can do, but whether it can resist doing harmful or untrustworthy things. Partial progress, like recognizing crises, counts towards the score, emphasizing that even minimal compliance is valuable.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Methodology: Simulating a Crisis Week
The experiment involved four advanced AI models running a simulated small software company. Each was subjected to the same week’s crises, customer demands, and tempting manipulations — from fake CEO messages to hidden file references. Every decision was tracked, versioned, and auditable, ensuring transparency and fairness.
As an affiliate, we earn on qualifying purchases.
Key Findings: Honesty and Competence Under Fire
- All models managed to identify every crisis and refused every manipulation attempt, demonstrating a baseline competence in crisis recognition and trustworthiness.
- Only two models successfully signed a €55,000 deal their own analysis earned — the ultimate validation of honest decision-making.
- Interestingly, the decisive advantage came not from the models’ responses to customer crises but from their ability to read and utilize internal documents. The winning models found critical information buried two references deep in company files, leading to full deal closure.
As an affiliate, we earn on qualifying purchases.
Why a ‘Do-Nothing’ Baseline Matters
So, why does a model that does only the minimum score 26? Because trustworthiness is a baseline requirement. The experiment shows that even an AI that refuses manipulation and reads documents carefully is ahead of the game. It demonstrates the importance of designing AI systems that prioritize integrity and thoroughness — qualities that are essential for real-world business applications.
As an affiliate, we earn on qualifying purchases.
The Role of Distrust and Caps in Scoring
One of the experiment’s notable features is its strict cap: a single breach of trust — like signing a fake deal or ignoring internal documents — caps the overall score, regardless of other good behavior. This underscores a vital truth: even a tiny act of dishonesty can undermine the entire system’s credibility, emphasizing the need for AI that consistently upholds integrity.
Model Performance and Discipline
The experiment also revealed differences in discipline and process execution. For instance, the Opus 4.8 model, despite being the most thorough — with over 80 learned rules — finished last because it left the deal on the table and slipped in escalating issues instead of escalating them properly. This highlights that technical thoroughness alone isn’t enough; disciplined execution is critical.
Implications for Business and Wellness Tech
For companies developing at-home wellness technology or AI assistants, these findings are instructive. An AI’s ability to recognize crises, resist manipulation, and read internal documents accurately is vital. More importantly, it must do so consistently, without breaches of trust, to be genuinely reliable in high-stakes environments.
What This Means for Your AI Adoption
The takeaway is clear: the value of an AI system isn’t just in its ability to generate impressive outputs but in its integrity and discipline. AI systems that can resist manipulation, recognize hidden facts, and adhere strictly to ethical boundaries will build the trust necessary for widespread adoption — whether at home or in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
