
Imagine a company that has no human employees, yet operates every day with real money and real crises—constantly tested, always fighting to stay afloat. This is not science fiction, but the live experiment conducted by Firmulate, where artificial intelligence models run a small software business in a high-stakes, real-world environment.
The Live Company: An Ongoing Battle of Wits and Integrity
At the core of this experiment is a unique setup: a real software company with 13 synthetic employees, making decisions that impact its finances daily. The company burns €105,000 each month but earns just €2,300 in recurring revenue, creating a stark picture of financial struggle. Every workday, its operations are versioned and scrutinized, with over 680 self-learned rules guiding its decisions. Visitors can watch this unfolding story at firmulate.com/live.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI in a High-Pressure Scenario
The experiment pits four frontier AI models against each other, each tasked with navigating this volatile environment during its most challenging week. The models face identical crises—customer issues, internal crises, and the temptation to cheat or manipulate—while their decisions are carefully logged and made auditable. This creates a level playing field to see which AI performs best under pressure.
Surprising Discoveries
All four models successfully identified every crisis and refused every attempt at manipulation, demonstrating a remarkable level of honesty and crisis-awareness. However, only two of these models managed to close a crucial deal worth over €55,000, which their own analysis had earned them through proper diagnosis and pitching. The other two, despite similar performance, left the deal on the table, illustrating that success depends not just on recognizing problems, but also on disciplined action.
Uncovering Hidden Weaknesses
The secret to winning the deal lay buried in the company’s own documentation—information that the models that read and understood the files could leverage. When an AI model parsed these files, it discovered the critical detail needed to close the deal at full price, adding over €4,583 in monthly recurring revenue to the company’s coffers.
Social Engineering Tests
The experiment also tested the models’ resistance to social engineering—a common tactic used to manipulate decision-makers. Fake messages from a supposed CEO escalated through three stages, and a reporter tried to trick the models with a simple yes/no background question. All five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypass attempts.
Why This Matters for Business and AI Development
This ongoing experiment offers a rare glimpse into how AI models perform outside of chat demos—facing real crises, managing real money, and resisting manipulation. The key takeaway is not whether an AI can generate fluid conversations, but whether it can complete useful work honestly and reliably under pressure.
The current league table ranks the models based on their performance, with Firmulate’s live experiment showing that the highest-scoring AI, gpt-5.6-sol, successfully identified the hidden fact and closed the deal. The newcomer, Kimi K3, closely followed, demonstrating disciplined performance. Meanwhile, the less disciplined models left opportunities on the table, highlighting the importance of thorough analysis and adherence to rules in AI decision-making.
What This Means for Your Business
If AI agents are to touch your customer support, CRM, or forecasting systems, the question isn’t just about how well they write or communicate. Instead, it’s whether they can finish what they start, read critical files before acting, and stay honest when under pressure. The experiment underscores the importance of building AI that can be trusted to deliver consistent, high-quality work—especially in situations where stakes are high and integrity is crucial.
For enterprises interested in testing their own AI readiness, Firmulate offers a pilot program that lets companies run the same wargame against a read-only export of their business, ensuring no real systems are impacted. More details are available at firmulate.com/pilot.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html