Control Safety · Flagship Result
Learning When to Act
Communication-efficient RL that jointly learns control and sampling time under a Lyapunov run-time assurance shield.
Abstract
Deep reinforcement learning produces capable controllers, but classical implementations update control and communicate at every time step. That is wasteful for bandwidth-limited platforms like high-altitude balloons, and it comes with no safety guarantee. This work asks: can a policy learn not just what to do, but when acting is worth it?
We design a communication-efficient reinforcement learning framework that jointly learns the control action and the inter-sample time, wrapped in a Lyapunov run-time assurance (RTA) shield. When the learned policy's proposed action would leave the certified region, a backup controller takes over, so efficiency is gained without giving up the safety certificate. The result: 3.51× fewer control updates than classical methods, with safety maintained throughout training and deployment.
A companion line of work inverts the setting: if a controller reveals when it acts, an adversary can learn sparse denial-of-service attacks against exactly those moments. That adversarial RL study is in preparation for ACC 2027.
Adam Haroon