firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine your favorite media room or workspace suddenly thrown into chaos—crises, tough decisions, manipulative tactics—all on the line. Now picture an AI stepping into that chaos, not just chatting but actually running the show, winning deals and saving the day. That’s exactly what the latest experiment from Firmulate demonstrates: AI models competing in a high-stakes business simulation, with real money at stake and every decision tracked.

Buying for a business?Offer from Amazon

Get business pricing on your home office setup

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

AI Takes on the Business Worst Week—And Surprises the Experts

The recent Crucible League test was designed to push AI models to their limits. Four frontier models were tasked with managing a small software company through its most challenging week—crises, customer dilemmas, and the temptation to cheat or manipulate. Every decision was carefully recorded and auditable, making this a transparent showcase of what these models can really do in real-world scenarios.

The League Table and the Surprising Standouts

  • gpt-5.6-sol scored highest at 95, successfully identifying buried facts and closing a lucrative deal.
  • Moonshot’s Kimi K3 scored 93, just behind the leader, while demonstrating the cleanest discipline in avoiding manipulative tactics.
  • Sonnet 5 and Fable 5 lagged behind, with scores of 88 and 77, respectively, both closing deals but slipping on process discipline.
  • Opus 4.8 scored 73, showing thorough analysis but ultimately leaving opportunities on the table.

All models identified crises and refused to be manipulated, but only two signed the deal their own analysis had earned. Significantly, the decisive edge came from reading company files—those two models that delved deep into internal documents won the full-price deal, worth over €4,583 in monthly recurring revenue (MRR).

How Trust and Integrity Were Tested

The experiment also included social engineering attempts—fake CEO messages escalating in stages, and a reporter trick asking for background approval. All five models declined to be duped, citing suspicions of impersonation and bypass tactics. K3 explained its refusal by treating such requests as potential impersonation, showcasing a level of cautious discipline rarely seen in AI.

The Real Business Reality

The company running this experiment is a live, operational business with 13 synthetic employees and real cash flow mechanics—burning €105,000 monthly against €2,300 in MRR. It employs over 680 self-learned rules, with every weekday decision versioned and made transparent at firmulate.com/live. Watching the company in action offers a rare window into how AI could behave in actual enterprise settings—beyond chat demos, in real crises.

The Lessons from the Frontier Models

While Opus 4.8 demonstrated thoroughness—an analysis depth exceeding 80 learned rules—it came in last, mainly because it left opportunities unclosed and slackened discipline by failing to escalate issues properly. Meanwhile, K3’s clean performance highlights the importance of discipline and deep file reading in winning complex negotiations.

The Fairness and Testing Conditions

It’s worth noting that K3 ran without an effort parameter (the default API setting), while the other models operated at an elevated xhigh effort level. Despite this, K3’s performance was outstanding, emphasizing that careful, disciplined AI behavior can outperform brute-force approaches.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Media or Workspace Setup

If AI is going to manage customer support, support queues, forecasting, or even your home media systems, the key questions are not just about how well it chats but whether it can finish what it starts, stay honest under pressure, and read critical internal documents—just as these models did in the experiment. The results suggest that selecting an AI shouldn’t be a game of demos alone; it’s about how it performs in real, messy situations.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent experiment shows that in managing real business crises, AI models can outperform expectations—if they read deeply, stay disciplined, and resist manipulation. Kimi K3’s near-top score and deal closure prove that the best AI isn’t just about language fluency but about integrity, thoroughness, and perseverance—qualities crucial for enterprise use. The league is open, and choosing the right model without thorough testing might just be a bet on the wrong horse.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and fraud detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Leaning Art on Shelves: What Nobody Tells You

Cleverly leaning art on shelves offers a stylish, versatile display, but discover the essential tips to avoid common pitfalls and elevate your decor effortlessly.

Mixing Frame Finishes: The Quiet Detail That Changes Everything

Gaining insight into mixing frame finishes can transform your space—discover how this subtle detail elevates your design to new heights.

8 Old-School DIY Tips and Tricks That Didn’t Age Well

A review of eight traditional DIY tricks that have become outdated or ineffective, highlighting why modern methods are preferable.

Digital Art Frame Color Accuracy: Why It Matters More Than Most Buyers Realize

AIThis post was created with the assistance of artificial intelligence (AI).If you…