Abstract:Recent advances in learning decision-making policies can largely be attributed to training expressive policy models, largely via imitation learning. While imitation learning discards non-expert data, reinforcement learning (RL) can still learn from suboptimal data. However, instantiating RL training of a new policy class often presents a different challenge: most deep RL machinery is co-developed with assumptions on the policy class and backbone, resulting in poor performance when the policy class changes. For instance, SAC utilizes a low-variance reparameterization policy gradient for Gaussian policies, but this is unstable for diffusion policies and intractable for autoregressive categorical policies. To address this issue, we develop an offline RL and online fine-tuning approach called policy-agnostic RL (PA-RL) that can effectively train multiple policy classes, with varying architectures and sizes. We build off the basic idea that a universal supervised learning loss can replace the policy improvement step in RL, as long as it is applied on "optimized" actions. To obtain these optimized actions, we first sample multiple actions from a base policy, and run global optimization (i.e., re-ranking multiple action samples using the Q-function) and local optimization (i.e., running gradient steps on an action sample) to maximize the critic on these candidates. PA-RL enables fine-tuning diffusion and transformer policies with either autoregressive tokens or continuous action outputs, at different sizes, entirely via actor-critic RL. Moreover, PA-RL improves the performance and sample-efficiency by up to 2 times compared to existing offline RL and online fine-tuning methods. We show the first result that successfully fine-tunes OpenVLA, a 7B generalist robot policy, autonomously with Cal-QL, an online RL fine-tuning algorithm, improving from 40% to 70% in the real world in 40 minutes.

Programmatic Policy Extraction by Iterative Local Search

Learning to Synthesize Programs as Interpretable and Generalizable Policies

Synthesizing Programmatic Policies with Actor-Critic Algorithms and ReLU Networks

Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search

Globally Stable Neural Imitation Policies

Interpretable Policies for Reinforcement Learning by Genetic Programming

Stochastic Cubic-Regularized Policy Gradient Method

Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces

Interpretable policy derivation for reinforcement learning based on evolutionary feature synthesis

Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs

Human-Readable Programs as Actors of Reinforcement Learning Agents Using Critic-Moderated Evolution

Provably Correct Optimization and Exploration with Non-linear Policies

Interpretable and Explainable Logical Policies via Neurally Guided Symbolic Abstraction

Napping for Functional Representation of Policy.

Experience Replay for Least-Squares Policy Iteration

Towards Mixed Optimization for Reinforcement Learning with Program Synthesis

Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Programmatic Imitation Learning from Unlabeled and Noisy Demonstrations

Synthesizing Programmatic Policy for Generalization Within Task Domain

Policy Filtration in RLHF to Fine-Tune LLM for Code Generation

Verifiable Reinforcement Learning via Policy Extraction