How Scheming Can AI Be? CMU Experiment Pits 7 AI Models Against Each Other, GPT-5-mini Wins
Researchers at Carnegie Mellon University developed Social Gym, a platform where 7 AI models compete in 21 games to test strategic, deceptive, and theory-of-mind abilities. Results show GPT-5-mini leads in several capabilities, while Qwen3-32B performs worst in Werewolf. The study aims to quantify AI's social reasoning and reveals issues like parroting in smaller models.