Scale X released a browser game that simulates human-in-the-loop approval of AI coding agents, where players decide to approve or reject commands under time pressure. The game collected data from over 40,000 runs and 409,000 individual approve/reject decisions.
The core finding is that the average player missed one-third of threats, with an average accuracy of 66.3%. Additionally, 32.9% of sessions ended with a negative score, meaning penalties for approving threats and blocking safe commands outweighed correct actions.
Notably, 35.2% of players caught all threats, but only 20.8% of them did so while blocking no more than one-fifth of safe commands; the rest may have achieved high scores by over-blocking. Another 7% of players approved all commands, jokingly called 'fans of dangerous skip permissions.'