SSignal89
AI前线
1 sourcesMajor AI Paper: Distillation Defenses of Top Three Global Models Fully Broken—Small Models Extract Hidden Reasoning Chains from Large Models, Kimi-K3 Shows Anomalous Reproduction Probability
A new study reveals severe security vulnerabilities in the APIs of Anthropic, OpenAI, and Google, allowing attackers to use lightweight models as "decoders" to recover hidden reasoning chains from encrypted reasoning blocks. Led by MATS program researcher Alexander Panfilov and co-authored by the Max Planck Institute for Intelligent Systems and others, the paper has garnered over 2.1 million reads. Additionally, the study found that Kimi-K3 reproduces Claude/GPT reasoning fragments with anomalously high probability, though the authors stress this is not sufficient to prove distillation.