Closing the Loop · Safe Learning
Safe Preference-Based RL
Learning policies offline from preferences, labeled by vision-language models, without trading away safety.
Abstract
Reward engineering is brittle; preferences are a more natural supervision signal, and vision-language models can now provide them at scale. But preference-based RL, like all RL, will happily trade safety for preference satisfaction unless safety is built in as a constraint.
This project develops state-conditioned safe offline preference-based reinforcement learning: policies learned entirely offline from preference labels, with safety constraints conditioned on state, so the learned policy inherits both what humans (or VLMs) prefer and what safety demands, without online trial and error on a physical system.
It is the control-side half of the closed loop this research program is building: VLM-aligned perception (LCLA) on one end, safely-learned policies from VLM supervision on the other.
Adam Haroon