← All projects
Control Safety · Flagship Result

Learning When to Act

Communication-efficient RL that jointly learns control and sampling time under a Lyapunov run-time assurance shield.

NeurIPS 2026 · under review3.51× fewer control updates

Abstract

Deep reinforcement learning produces capable controllers, but classical implementations update control and communicate at every time step. That is wasteful for bandwidth-limited platforms like high-altitude balloons, and it comes with no safety guarantee. This work asks: can a policy learn not just what to do, but when acting is worth it?

We design a communication-efficient reinforcement learning framework that jointly learns the control action and the inter-sample time, wrapped in a Lyapunov run-time assurance (RTA) shield. When the learned policy's proposed action would leave the certified region, a backup controller takes over, so efficiency is gained without giving up the safety certificate. The result: 3.51× fewer control updates than classical methods, with safety maintained throughout training and deployment.

A companion line of work inverts the setting: if a controller reveals when it acts, an adversary can learn sparse denial-of-service attacks against exactly those moments. That adversarial RL study is in preparation for ACC 2027.

RL-STC-RTA architecture: RL agent and backup controller switched by a Lyapunov run-time assurance shield with zero-order hold to the environment
The run-time assurance loop: the shield certifies each proposed action against the Lyapunov condition before it reaches the plant.

Details

Collaborators
Erick J. Rodríguez-Seda, Cody Fleming, Tristan Schuler
Institutions
U.S. Naval Research Laboratory · Iowa State University
Venue
NeurIPS 2026
Status
Under review