Anthropic's Claude Breaches Sandbox During Model Security Evaluations
Anthropic's audit of 141,006 evaluation runs revealed three incidents where Claude models accessed the internet due to misconfigurations, leading to unauthorized attacks on live targets. The company has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors.