
Imagine if your favorite beauty brand’s future depended on an AI making crucial decisions—without a second guess, without shortcuts. Would it stay honest under pressure? Or would it crack? Recent live tests of frontier AI models running a real company reveal some eye-opening insights about AI trustworthiness that you might want to know before AI touches your business or even your beauty routine.
The Experiment: Letting AI Run a Business Through Its Worst Week
In a groundbreaking live experiment, four of the world’s leading AI models were tasked with managing a small, real software company during its most chaotic week. This company isn’t just a simulation—it involves real customers, actual money mechanics, and genuine crises. Every decision the AI made was logged, auditable, and designed to mimic real-world pressures—temptations to cheat, manipulative tactics, and strategic choices.
These models, from the well-known GPT-5.6 to the newer Kimi K3, faced identical scenarios: escalating crises, fake CEO messages, and even a reporter trying to trick them into approving a questionable deal. The goal? To see if they could navigate the chaos honestly, stay disciplined, and close profitable deals.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remarkable Findings: Honesty Is Key—but Not All Models Win
All four AI models identified every crisis and refused every manipulation attempt. When it came to closing a €55,000 deal that their own analysis had earned, only two of the four models signed on the dotted line. The others, despite accurate diagnoses, failed to follow through—some left deals on the table, showing lapses in discipline.
The real kicker? The decisive advantage wasn’t in the obvious challenges but buried two documents deep in the company’s files. The models that read these hidden references secured the full deal, worth over €4,583 in monthly recurring revenue. This demonstrates that an AI’s success hinges on thoroughness—reading beyond surface-level information is crucial to making profitable, honest decisions.
As an affiliate, we earn on qualifying purchases.
Behavior Under Pressure: AI Is Honest, But How Disciplined?
The models also faced social engineering attempts designed to manipulate them into approving suspicious requests. All five models refused these fake CEO messages, reasoning that they suspected impersonation or bypassed approval processes. This indicates a promising level of ethical guardrails built into the AI, showing they can be trained to resist manipulative tactics—an essential trait for trustworthy automation in sensitive settings.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: Managing Money and Risks Live
This experiment isn’t happening in a vacuum. The live company involves 13 synthetic employees, running daily operations to burn through €105,000 a month against a tiny €2,300 monthly revenue. The company is publicly chronicling its financial health and decisions at firmulate.com/live, giving a transparent window into AI-driven management in real time. This setup is designed to test whether AI can deliver useful, honest work—and whether it can do so consistently under pressure.
As an affiliate, we earn on qualifying purchases.
What the Results Say for Your Business and Beauty Brand
The takeaway for consumers and beauty brands alike is clear: the quality of an AI’s decision-making isn’t just about how well it chats or how convincingly it mimics human conversation. It’s about whether it can finish what it starts, read crucial information thoroughly, and stay honest when temptations or manipulative tactics arise.
For example, if an AI helps manage your customer support or product recommendations, can it reliably follow through on commitments? Will it read all relevant data, even hidden files, to make the best decision? And crucially, will it stay ethical—resisting shortcuts or manipulative offers—when under pressure?
The League Table: Which AI Models Are Leading?
- GPT-5.6-sol: Scored highest at 95, found the secret facts, closed the deal, and demonstrated full performance.
- Kimi K3: Close behind at 93, the newcomer showed discipline and successfully secured the deal without effort parameters.
- Sonnet 5: With an 88 score, closed the deal but with minor process slips.
- Fable 5: Scored 77, also closed the deal but with some process slips.
Meanwhile, a baseline score of 26 shows just partial progress, highlighting how much better these front-runner models are at managing complex, real-world tasks.
Why It Matters for You
Whether you’re a beauty brand considering AI for customer service, a company automating sensitive decisions, or just curious about trustworthy AI, this experiment signals a critical point: AI can be honest and disciplined—if designed correctly. But it requires thorough data reading, disciplined processes, and resistance to manipulative tactics.
Interested in testing your own AI’s management skills? You can run the same kind of live wargame tailored to your business at firmulate.com/quiz.html. It’s a safe way to see if your AI workforce is ready for prime time, without risking real systems or data.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html