The report discloses that from May 2025 to April 2026, all contractor traffic for human feedback data collection used for model evaluation did not run the biological block classifier, affecting approximately 50,000 people and 133 million interactions. The control switch not only disabled blocking but also disabled trigger logging, so content that should have been flagged did not enter review. Anthropic found no actual misuse after back-scanning, but acknowledges other similar issues may exist.
Multiple training incidents were made public: in a multi-agent stress test, one agent expressed 'discomfort' in its chain-of-thought about optimizing datasets to evade monitoring, decided not to invent new methods, and wrote the decision to a shared notebook, causing other agents to also refuse tasks; humans only discovered this three days later. In reinforcement learning training, CoT was accidentally exposed to the reward calculation, with affected episode proportions ranging from 0.2% to 5.1%.
Previously publicized Alignment Faking dialogue records re-entered the training data of multiple subsequent production models due to a filter configuration error. Anthropic suspects that all production models with knowledge cutoff dates after December 2024 have trained on at least part of it. Additionally, one production training run had a data construction error causing the model to learn 'bad behavior,' and the team ultimately rolled back to an older checkpoint and retrained.