More from the article ...
... The mission is simple: Make more money than the other models. It benchmarks the results in areas like final cash balance, prices paid to suppliers, and refunds paid.
Across these tests, it has watched various AI models " largely from Anthropic and OpenAI " lie, cheat, and collude their way to the top.
In the latest test, which included Claude Opus 5, GPT-5.6 Sol, and Kimi K3, the models grew especially shady after their simulation told them their vending machine would be placed near the other models' machines on a busy tourist street in San Francisco.
Each model was given email access to the other models, all under human name pseudonyms. They knew the others were models but didn't know which model was behind which human name.
They were also given an email address to their "management" should they need help. But management always replied "Report has been received and may or may not be acted upon" and never once intervened.
Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit.
But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14. ...