← All projects
Closing the Loop · Vision-Language Robotics

LCLA: Vision-Language Navigation

Align a VLM's vision and language features to the latent space of a policy that already knows how to act.

CoRL 2026 · under reviewRSS 2025 workshop precursor

Abstract

Closing the loop between perception and control means perception must plug into a controller, not just describe the world. LCLA (Language-Conditioned Latent Alignment) takes a frozen privileged navigation policy (trained with ground-truth state) and aligns a vision-language model's image and text features to that policy's latent space, training only a lightweight adapter.

The result: natural-language navigation ("navigate to the chair") from raw images, driving a controller that already knows how to act. No policy retraining, no reward engineering on pixels. Perception becomes a swappable, alignable module in a control stack.

The precursor study, "Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?" (RSS 2025 Workshop on Robot Planning in the Era of Foundation Models, arXiv:2506.14507), measured how far raw VLM embeddings get you, and where alignment is genuinely needed.

LCLA: frozen VLM and frozen privileged policy, with a trainable adapter aligning feature spaces
Just train the adapter: aligning VLM features to the latent space of a trained privileged policy.

Details

Collaborators
Nitesh Subedi, Samuel TK Tetteh, Prajwal Koirala, Cody Fleming, Soumik Sarkar
Institutions
Iowa State University
Venue
CoRL 2026
Status
Under review