Back to feed
News Story
APriority72
Hacker News (AI filter)
1 sources

Humans Missed 1 in 3 Threats When Approving AI Agent Commands

A study across 40,000 game runs found that humans missed one in three threats when approving AI agent commands, highlighting the risks of human oversight in AI agent deployment. The findings suggest that relying on human approval may not be sufficient for safety, and more robust safeguards are needed.

SynthePulse Insight · AI deep readingMembers

Human Oversight of AI Agents Has Blind Spots: 40,000 Game Runs Reveal One-Third Threat Miss Rate

Version 1 · 1 source

A study based on 40,000 game runs shows that humans miss an average of one-third of threats when approving AI agent commands, and the most dangerous credential-stealing commands are missed three times as often as clearly destructive ones.

  • Average players missed one-third of threats, with an accuracy of only 66.3%.
  • 32.9% of sessions ended with a negative score, where penalties outweighed correct actions.
  • 35.2% of players caught all threats, but only 20.8% did so while blocking no more than one-fifth of safe commands.
Open section navigationResearch Background and Core Data

Research Background and Core Data

Scale X released a browser game that simulates human-in-the-loop approval of AI coding agents, where players decide to approve or reject commands under time pressure. The game collected data from over 40,000 runs and 409,000 individual approve/reject decisions.

The core finding is that the average player missed one-third of threats, with an average accuracy of 66.3%. Additionally, 32.9% of sessions ended with a negative score, meaning penalties for approving threats and blocking safe commands outweighed correct actions.

Notably, 35.2% of players caught all threats, but only 20.8% of them did so while blocking no more than one-fifth of safe commands; the rest may have achieved high scores by over-blocking. Another 7% of players approved all commands, jokingly called 'fans of dangerous skip permissions.'

Free for now

Read the full analysis

5 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This report is based on game data released by Scale X, sourced from game run records, but the game environment differs from real scenarios (e.g., higher threat ratio, users aware of being tested). Some conclusions (like fatigue effects) reference Anthropic's observations, but the game data itself is the source claim.

Primary report

Hacker News (AI filter)

Primary source