Closing the Loop · Vision-Language Robotics
LCLA: Vision-Language Navigation
Align a VLM's vision and language features to the latent space of a policy that already knows how to act.
Abstract
Closing the loop between perception and control means perception must plug into a controller, not just describe the world. LCLA (Language-Conditioned Latent Alignment) takes a frozen privileged navigation policy (trained with ground-truth state) and aligns a vision-language model's image and text features to that policy's latent space, training only a lightweight adapter.
The result: natural-language navigation ("navigate to the chair") from raw images, driving a controller that already knows how to act. No policy retraining, no reward engineering on pixels. Perception becomes a swappable, alignable module in a control stack.
The precursor study, "Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?" (RSS 2025 Workshop on Robot Planning in the Era of Foundation Models, arXiv:2506.14507), measured how far raw VLM embeddings get you, and where alignment is genuinely needed.
Adam Haroon