← All projects
Closing the Loop · Safe Learning

Safe Preference-Based RL

Learning policies offline from preferences, labeled by vision-language models, without trading away safety.

ICLR 2027 · in preparation

Abstract

Reward engineering is brittle; preferences are a more natural supervision signal, and vision-language models can now provide them at scale. But preference-based RL, like all RL, will happily trade safety for preference satisfaction unless safety is built in as a constraint.

This project develops state-conditioned safe offline preference-based reinforcement learning: policies learned entirely offline from preference labels, with safety constraints conditioned on state, so the learned policy inherits both what humans (or VLMs) prefer and what safety demands, without online trial and error on a physical system.

It is the control-side half of the closed loop this research program is building: VLM-aligned perception (LCLA) on one end, safely-learned policies from VLM supervision on the other.

Details

Collaborators
Cody Fleming
Institutions
Iowa State University · VRAC
Venue
ICLR 2027
Status
In preparation

Links

Status
Manuscript in preparation
Related
LCLA · Learning When to Act
Contact
aharoon@iastate.edu