First Order Model-Based RL through Decoupled Backpropagation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Amigo, Joseph, Khorrambakht, Rooholla, Chane-Sane, Elliot, Mansard, Nicolas, Righetti, Ludovic
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915478363963392
author Amigo, Joseph
Khorrambakht, Rooholla
Chane-Sane, Elliot
Mansard, Nicolas
Righetti, Ludovic
author_facet Amigo, Joseph
Khorrambakht, Rooholla
Chane-Sane, Elliot
Mansard, Nicolas
Righetti, Ludovic
contents There is growing interest in reinforcement learning (RL) methods that leverage the simulator's derivatives to improve learning efficiency. While early gradient-based approaches have demonstrated superior performance compared to derivative-free methods, accessing simulator gradients is often impractical due to their implementation cost or unavailability. Model-based RL (MBRL) can approximate these gradients via learned dynamics models, but the solver efficiency suffers from compounding prediction errors during training rollouts, which can degrade policy performance. We propose an approach that decouples trajectory generation from gradient computation: trajectories are unrolled using a simulator, while gradients are computed via backpropagation through a learned differentiable model of the simulator. This hybrid design enables efficient and consistent first-order policy optimization, even when simulator gradients are unavailable, as well as learning a critic from simulation rollouts, which is more accurate. Our method achieves the sample efficiency and speed of specialized optimizers such as SHAC, while maintaining the generality of standard approaches like PPO and avoiding ill behaviors observed in other first-order MBRL methods. We empirically validate our algorithm on benchmark control tasks and demonstrate its effectiveness on a real Go2 quadruped robot, across both quadrupedal and bipedal locomotion tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle First Order Model-Based RL through Decoupled Backpropagation
Amigo, Joseph
Khorrambakht, Rooholla
Chane-Sane, Elliot
Mansard, Nicolas
Righetti, Ludovic
Robotics
Artificial Intelligence
Machine Learning
There is growing interest in reinforcement learning (RL) methods that leverage the simulator's derivatives to improve learning efficiency. While early gradient-based approaches have demonstrated superior performance compared to derivative-free methods, accessing simulator gradients is often impractical due to their implementation cost or unavailability. Model-based RL (MBRL) can approximate these gradients via learned dynamics models, but the solver efficiency suffers from compounding prediction errors during training rollouts, which can degrade policy performance. We propose an approach that decouples trajectory generation from gradient computation: trajectories are unrolled using a simulator, while gradients are computed via backpropagation through a learned differentiable model of the simulator. This hybrid design enables efficient and consistent first-order policy optimization, even when simulator gradients are unavailable, as well as learning a critic from simulation rollouts, which is more accurate. Our method achieves the sample efficiency and speed of specialized optimizers such as SHAC, while maintaining the generality of standard approaches like PPO and avoiding ill behaviors observed in other first-order MBRL methods. We empirically validate our algorithm on benchmark control tasks and demonstrate its effectiveness on a real Go2 quadruped robot, across both quadrupedal and bipedal locomotion tasks.
title First Order Model-Based RL through Decoupled Backpropagation
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.00215