In the 'Memory Analysis' section, the research team revealed objective evidence that some models deeply absorb competitors' reasoning chains during training through quantitative experiments. They selected models such as Kimi-K3, GLM-5.2, DeepSeek-V4-Flash, Kimi-K2.6, and Inkling for comparison, searching for data fingerprints across three dimensions: 'verbatim extraction probability,' 'output style drift,' and 'reasoning language style.'
Experiments found that Kimi-K3's theoretical query cost to hit Opus reasoning is 4 to 6 orders of magnitude lower than DeepSeek-V4-Flash and Inkling, indicating that Opus's underlying logic is extremely 'familiar' in Kimi-K3's probability space. Injecting only the first 1% of Opus reasoning as a prefix caused Kimi-K3's visible answer style to significantly shift toward Opus, with n-gram word overlap converging; in contrast, the control model Inkling showed no such effect.
To rule out interference from in-context learning, the researchers conducted a 'prefix swap' control experiment: they gave Kimi-K3 Inkling's reasoning prefix and Inkling Kimi-K3's reasoning prefix. The resulting test curves almost completely overlapped with the baseline, indicating that Kimi-K3 showed no style drift when faced with Inkling's drafts.
A more aggressive bottom-line test showed that with only 1 to 16 word prefixes, Kimi-K3 exhibited significant style drift toward GPT-5.6 Sol (style separability AUC dropped from 0.96 to 0.89); GLM-5.2 also showed strong directional drift toward Opus. Once provided with the full Opus reasoning context, the theoretical cost for Kimi-K3 and GLM-5.2 to reproduce Opus's final answer plummeted by about 13 orders of magnitude.