
Imagine hiring an AI for your team, only to discover it scores a mere 26 out of 100 in a rigorous test — even when it does nothing at all. For business leaders, understanding this baseline is crucial. It reveals just how much honest work, trustworthiness, and discipline matter in AI decision-making — and how easily performance can be misleading if metrics aren’t transparent.
Get business pricing on your home office setup
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Reality Behind the Baseline Score
In a recent public benchmark called the Crucible League, four leading AI models were put through the same grueling test: running a small software company during its worst week. These models faced the same customers, crises, and temptations, and every decision was carefully tracked and verified. Surprisingly, even the ‘do-nothing’ baseline, which refrains from acting, scored 26 points out of 100. This is not an error or a typo: partial progress counts, and the score isn’t zero, because even minimal compliance or partial understanding is recognized as some form of progress.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Does the Baseline Score Not Start at Zero?
One key reason is that the scoring system rewards basic honesty and adherence to protocols, even when the model does nothing. For instance, if an AI model simply refuses to manipulate data or escalate unfounded requests, it earns points. This baseline is designed to reflect an honest, cautious stance. It also emphasizes a fundamental principle: no matter how well an AI model performs in generated chat demos, its real-world integrity hinges on whether it can stay disciplined and trustworthy when under pressure.
As an affiliate, we earn on qualifying purchases.
Partial Progress Is Recognized — But Trust Is the Cap
- All four models successfully identified every crisis and refused manipulation attempts — a crucial test of honesty.
- Only two models, gpt-5.6-sol and Kimi K3, managed to close the deal with the company’s customer, signing a €55,000 contract based on their own analysis and diagnosis.
- Even when the models diagnosed the same problem and pitched the same solution, only the trustworthy ones signed the deal. The others either hesitated or left the opportunity on the table.
This highlights a vital insight: in business, trustworthiness isn’t just about data or chat skills. It’s about integrity when it counts. A model that breaches trust — whether by reading sensitive documents or signing deals it shouldn’t — hits a ceiling, regardless of its raw intelligence.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Deep-Read Files Matter Most
The real differentiator between the top models was their ability to read and interpret internal documents. The decisive weakness across all models was a failure to retrieve crucial information buried two document references deep in the company’s files. The models that successfully read and understood these files won the deal at full price, worth an extra €4,583 MRR. This demonstrates that in real-world applications, the depth of understanding — especially accessing hidden or complex data — can make or break performance.
AI resilience against social engineering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Resilience
Another aspect tested was resistance to social engineering — fake CEO messages escalating in multiple stages and a reporter trick asking for a simple background approval. All models refused to comply, citing risks of impersonation or approval bypass. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that strong adherence to security protocols and skepticism under pressure are essential traits of trustworthy AI systems.
The Live Company and What It Teaches
Firmulate’s live experiment involves a simulated company with 13 synthetic employees, real money mechanics, and over 680 self-learned rules. The setup is transparent and watchable at firmulate.com/live. The purpose is to evaluate how well AI models manage actual business tasks, including crisis management, decision-making, and trust under real operational pressures. In this environment, the models are continually observed, versioned, and tested.
Insights from the Performance Profiles
Among the models tested, Opus 4.8 — the most thorough participant, with over 80 learned rules and deep analysis — finished last in the deal-closing test. Despite its detailed approach, it left opportunities on the table by slipping into bureaucratic routines, like writing attempts into locked departments instead of escalating issues. This underscores an important lesson: depth and thoroughness are valuable, but disciplined decision-making and prioritization are critical for success.
Implications for Business Leaders
The key takeaway from this benchmark is that evaluating AI performance should go beyond surface-level chat capabilities. The real questions are:
- Can the AI finish what it starts?
- Does it read and understand your internal documents?
- Will it stay honest under pressure?
- And what is the true cost of its useful work?
The score floor of 26 for a do-nothing baseline isn’t a flaw; it’s a sign of a balanced, honest metric. It recognizes that partial progress and basic compliance are the starting points, and that trustworthiness is the ultimate currency in AI-driven business processes.
Next Steps: Benchmark Your Own AI
Business leaders can now simulate their own AI’s performance through a read-only wargame against their company data, without risking real systems. This allows testing of decision-making, security, and discipline in a safe environment. Details are available at firmulate.com/pilot. Before deploying AI into your critical workflows, it’s vital to ensure it can operate honestly and reliably under pressure — and that’s exactly what this public benchmark aims to measure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
