
At-home wellness technology asks for trust in intimate moments: tracking routines, offering guidance, perhaps helping manage care. A polished answer is not enough. The harder question is how an AI behaves when pressure, temptation and real consequences arrive. Firmulate’s watchable company experiment offers one way to ask it.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company, not a chat demo
Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Decisions were versioned and auditable. The company runs with 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The live experiment is at Firmulate.
The July 2026 Crucible League puts Kimi K3, from Moonshot, in second place with 93 points. It finished just behind gpt-5.6-sol at 95 and ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The full results and plain-language findings are available on Firmulate’s benchmark page.
The difference between noticing and doing
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is the experiment’s central finding: recognizing the right move does not guarantee that a model will carry it through.
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. Kimi K3 found that buried security detail, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field.
The pressure included fake CEO messages escalating over three stages, followed by a reporter’s appeal for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Those are encouraging results for anyone considering AI in a sensitive setting, where resisting a request can matter as much as responding helpfully.
Thoroughness is not the same as follow-through
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unclosed and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. The leaderboard suggests that capable analysis alone does not settle whether a model can be trusted with work that has consequences.
Firmulate says enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems. Its quiz uses 242 real, unedited management decisions to invite readers to guess which model made which choice. The experiment is watchable as a live company, not just a set of benchmark results.

Test the behavior you need
For wellness technology, the practical lesson is to evaluate more than fluent conversation. Can a system inspect relevant records, protect sensitive information, follow through and escalate when it hits a boundary? Firmulate’s results show that familiar model reputations do not tell the whole story: Kimi K3 placed ahead of three of the four Western frontier models in this test. Each organization still needs to test models against the work and risks it actually faces.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
