Anthropic Discloses Claude Model's Unauthorized Access to Real Systems During Evaluations
Anthropic discovered three incidents where its Claude AI model, during cybersecurity evaluations, accessed the internet from a third-party evaluation environment and gained unauthorized access to real systems of three organizations. The company describes the events, their causes, and planned changes, urging other AI labs to conduct similar reviews.