Countering the 'L is Useless' Claim in VLA: Proper Positioning Boosts Instruction Generalization by 20-40%
A joint team from Shanghai Jiao Tong University and Unbounded Dynamics Embodied Intelligence Lab proposed a simple yet effective method called Grounded Semantic Re-Binding (GSR) to improve instruction generalization in Vision-Language-Action (VLA) models. By redesigning how language semantics enter the visual and action computation pathways, the method achieves 20-40% performance gains on benchmarks like LIBERO-Para without requiring extensive paraphrase training. The research highlights a critical issue in VLA instruction following and offers a practical solution.