SSignal86
机器之心
1 sourcesRefCaptioner: Aligning Reference Images with Video Semantics
Peking University and Kling team propose RefCaptioner, which enhances the model's ability to recognize and bind reference images through post-training, and constructs structured video descriptions to address the decline in correspondence ability in multi-reference-image scenarios. The method uses mixed-data SFT and a dual-layer adaptive reward function HCD-GRPO to improve precise alignment between reference images and video semantics.