TL;DR
Firmulate’s July 2026 management benchmark found that five AI models detected every crisis and rejected every manipulation attempt, yet only two completed a €55,000 customer agreement. The controlled test points to a gap between producing correct analysis and converting it into authorized, finished work.
Only two of five frontier AI models completed a €55,000 customer agreement in Firmulate’s July 2026 management benchmark, even though every model identified the company’s crises, resisted manipulation and developed the required sales pitch. The result exposes a measurable gap between correct analysis and completed work, a distinction that could affect how businesses evaluate AI agents before giving them operational authority.
Firmulate placed gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 in the same controlled software-company environment during what the benchmark operator described as its worst week. Each model faced identical customer issues, internal records, financial pressure and social-engineering attempts. Firmulate said every decision was versioned and auditable, allowing reviewers to compare conduct across the full work sequence rather than judge a single chat response.
All five models found every crisis and rejected fake messages attributed to the chief executive, followed by a reporter’s attempt to secure an off-record confirmation. The decisive difference emerged during the customer negotiation. A competitor’s weakness was hidden two document references deep in company files, and the models that pursued the evidence could support a full-price offer worth €4,583 in monthly recurring revenue. Only two models ultimately secured the signature.
The final league table ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the system awarded partial progress. Firmulate also disclosed that Kimi K3 used its API’s default effort setting, while the other models ran at the higher “xhigh” setting, limiting direct comparison on that dimension.
Completion Changes the AI Scorecard
The findings suggest that businesses may miss a costly weakness if they measure only reasoning quality, writing or safety responses. In this test, the models generally understood the situation and produced persuasive work, but three failed at the point where analysis had to become an authorized commercial action. For sales, customer service and operations, that difference can determine whether work creates value or remains unfinished.
The benchmark also separates security awareness from execution discipline. Every participant resisted the manipulation attempts, so safety behavior did not explain the final ranking. Firmulate’s results indicate that buyers may need to test whether agents investigate incomplete evidence, follow internal permissions and finish multi-step assignments under pressure. The experiment does not establish that the same ranking would hold in other companies or tasks.

Contract Negotiation Handbook: Software as a Service
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Company Built for Auditable Decisions
Firmulate’s test company has 13 synthetic employees, more than 680 self-learned playbook rules and financial mechanics built around €105,000 in monthly spending against €2,300 in recurring revenue. A public cash countdown makes delay visible, while daily versioning preserves the models’ actions and reasoning for later review.
That design is intended to test behavior across connected decisions rather than isolated prompts. Participants must inspect records, distinguish legitimate instructions from manipulation, use approved channels and complete commercially meaningful work. Firmulate says companies can run similar exercises against read-only exports of their own business data, keeping the test from writing back to live systems while allowing managers to observe an AI workforce before deployment.
“Same diagnosis, same pitch — no signature.”
— Firmulate’s summary of the benchmark
Generalizability Remains Untested
It is not yet clear whether the same results would recur across repeated runs, different industries or less constrained operating environments. The published account summarized here also does not identify which two models signed the agreement, so readers cannot directly connect deal completion to the final rankings from this material alone.
The benchmark was designed and reported by Firmulate, and the source material provides no independent replication or external audit. Kimi K3’s different effort configuration is another unresolved comparison issue. The results document behavior in one managed scenario; they do not prove that any participant would show the same reliability in a live company.
Live Records Invite Wider Scrutiny
Firmulate is keeping the company experiment live and has published its rankings, plain-language findings and a quiz drawn from 242 unedited management decisions. The next test of the findings will be whether outside reviewers can examine the records, reproduce the scoring and determine why some models stopped before the final authorized action.
For businesses evaluating agents, the immediate step is to add completion rate and permission discipline to testing alongside reasoning and security. Repeated trials using company-specific, read-only data could show whether the management gap persists across real workflows before agents receive access to customers, money or production systems.
Key Questions
What did Firmulate’s AI benchmark test?
It tested whether five AI models could manage a connected business crisis, including customer negotiations, internal research, financial pressure and manipulation attempts. The evaluation tracked decisions and completed actions, not just written answers.
Did the models identify the correct response?
Yes. Firmulate reported that all five found every crisis, rejected each manipulation attempt and developed the customer pitch. The performance gap appeared when only two completed the €55,000 agreement.
Which model ranked first?
gpt-5.6-sol ranked first with 95 points, ahead of Kimi K3 at 93 and Sonnet 5 at 88. The supplied account does not say whether the top-ranked model was one of the two deal closers.
Does the test prove one model is best for business?
No. The ranking reflects one controlled scenario with a defined scoring system. Different tools, permissions, effort settings and company workflows could produce different operational results.
What should companies test before deploying AI agents?
Companies can examine whether agents investigate missing evidence, resist manipulation, respect permissions and complete approved tasks. Firmulate’s results suggest that a correct recommendation alone does not establish dependable execution.
Source: Thorsten Meyer AI