AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an employee who, even in the worst week, scores at least 26 out of 100—regardless of how badly things go. In the world of AI, that ’employee’ might be your next virtual assistant or decision-maker. But what does that baseline really tell us about how trustworthy and effective these AI models are? And why should your business care about the difference between a model that scores 26 and one that scores 95? The answers lie in a recent public experiment by Firmulate, which offers a transparent look into the true capabilities—and limitations—of leading AI models under real-world pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Challenge of Measuring AI Performance in Business Contexts

As AI becomes more integrated into business operations—whether managing customer relationships, support queues, or financial forecasts—it’s crucial to understand not just whether an AI can produce good-sounding answers, but whether it can be trusted to finish what it starts. Traditional benchmarks often focus on chat quality or language fluency, which can be misleading when assessing real-world utility. Firmulate’s recent experiment shifts the focus to decision-making under pressure, replicating the kinds of crises and temptations an AI might face when managing actual business processes.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Firmulate Experiment: A Real-World Test for AI Models

In this live setup, four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—were tasked with running a small software company through its worst week. This involved handling the same customers, same crises, and same opportunities for manipulation across all models, with every decision recorded and made auditable. The goal: see if the models can spot crises, refuse unethical requests, and ultimately secure a vital deal worth €55,000.

Amazon

business AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why 26 Points Is the Starting Line

When the AI models were evaluated, all managed to identify every crisis and refused all attempts at manipulation, such as fake CEO messages or backdoor requests. However, only two models—gpt-5.6-sol and Kimi K3—actually signed the deal that their own analysis suggested they could close. The other two, Sonnet 5 and Opus 4.8, left the deal on the table, despite having diagnosed the opportunity correctly. Interestingly, even a ‘do-nothing’ baseline, which essentially makes no decisions, scores 26 points. That means the lowest possible score for a model that attempts to act is not zero but 26, because partial progress counts toward the total, and any breach of trust caps the score.

Amazon

AI trustworthiness evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Single Breach Caps Performance

One key finding: a single instance of trust violation—the equivalent of an employee signing a deal they shouldn’t—reduces the overall score to the baseline floor of 26. This underscores a critical principle: in high-stakes decision environments, even small breaches can overshadow the good work the AI does elsewhere. It’s a reminder that trustworthiness is paramount, and that a model’s value isn’t just about spotting crises but also about integrity and discipline.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Files Matters

The most decisive advantage in the experiment was the model’s ability to read and interpret internal documents. The winning models found a buried reference two documents deep in the company’s files, which was crucial for closing the deal at full price—an extra €4,583 in monthly recurring revenue. This highlights a vital point for businesses: an AI’s capacity to access and interpret internal data can be the difference between a good decision and a missed opportunity.

Defense Against Manipulation and Social Engineering

All models successfully refused social engineering attempts, including staged fake CEO messages and reporter tricks. Kimi K3’s explanation was clear: it treats such requests as potential impersonation, refusing to act without explicit, verified approval. This demonstrates that well-designed AI models can be resilient against social engineering, a common threat in digital operations.

The Real Business Environment: A Live Company in Action

The experiment was run on a simulated but realistic company, with 13 synthetic employees using real money mechanics—burning €105,000 a month against €2,300 monthly recurring revenue. The environment was versioned daily, providing an ongoing, transparent test bed for evaluating AI decision-making, discipline, and honesty under real pressure. Visitors can watch this ongoing experiment at firmulate.com/live.

The Lessons for Business Leaders

This live benchmark demonstrates that AI performance isn’t just about generating appealing responses; it’s about how well models can recognize crises, uphold integrity, and complete their tasks reliably. A model that scores 95, like gpt-5.6-sol, not only finds critical information but also closes deals at full value—showing real business impact. Conversely, even the most thorough model, Opus 4.8, scored the lowest because it slipped on discipline, leaving opportunities on the table and escalating issues into locked departments instead of resolving them.

Why Trust in AI Matters More Than Ever

As AI models become embedded in your company’s decision processes, understanding their true capabilities and limitations is vital. The Firmulate live experiment offers a transparent, auditable way to assess whether an AI can be trusted under pressure, not just whether it can chat well. For business leaders, the key takeaway: don’t settle for high scores that don’t reflect real-world performance. Instead, look for models that can read deeply, refuse manipulation, and finish what they start—even if that means accepting a baseline score of 26 for do-nothing attempts.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Temperature-Controlled Mattresses and Pads Work

Offering personalized comfort through responsive heating and cooling, temperature-controlled mattresses and pads adapt to your body and environment, and here’s how they work.

Gpiozero Flow

Gpiozero Flow releases an updated version aimed at simplifying GPIO programming on Raspberry Pi, with new features and improved stability.

Destressing Headbands: How Wearable Meditation Aids Sleep

Pioneering wearable meditation headbands offer a new way to de-stress and improve sleep, but how exactly do they work?

Rapper Santy Sharma Makes Viral Joke And Meme About The New iPhone Duo’s Insane Price Tag

Rapper Santy Sharma’s joke about the new iPhone duo’s price has gone viral, sparking widespread online discussion and memes.