Accelerating Model-Based Reinforcement Learning using Non-Linear Trajectory Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Calì, Marco, Giacomuzzo, Giulio, Carli, Ruggero, Libera, Alberto Dalla
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908391551533056
author Calì, Marco
Giacomuzzo, Giulio
Carli, Ruggero
Libera, Alberto Dalla
author_facet Calì, Marco
Giacomuzzo, Giulio
Carli, Ruggero
Libera, Alberto Dalla
contents This paper addresses the slow policy optimization convergence of Monte Carlo Probabilistic Inference for Learning Control (MC-PILCO), a state-of-the-art model-based reinforcement learning (MBRL) algorithm, by integrating it with iterative Linear Quadratic Regulator (iLQR), a fast trajectory optimization method suitable for nonlinear systems. The proposed method, Exploration-Boosted MC-PILCO (EB-MC-PILCO), leverages iLQR to generate informative, exploratory trajectories and initialize the policy, significantly reducing the number of required optimization steps. Experiments on the cart-pole task demonstrate that EB-MC-PILCO accelerates convergence compared to standard MC-PILCO, achieving up to $\bm{45.9\%}$ reduction in execution time when both methods solve the task in four trials. EB-MC-PILCO also maintains a $\bm{100\%}$ success rate across trials while solving the task faster, even in cases where MC-PILCO converges in fewer iterations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Model-Based Reinforcement Learning using Non-Linear Trajectory Optimization
Calì, Marco
Giacomuzzo, Giulio
Carli, Ruggero
Libera, Alberto Dalla
Machine Learning
Robotics
This paper addresses the slow policy optimization convergence of Monte Carlo Probabilistic Inference for Learning Control (MC-PILCO), a state-of-the-art model-based reinforcement learning (MBRL) algorithm, by integrating it with iterative Linear Quadratic Regulator (iLQR), a fast trajectory optimization method suitable for nonlinear systems. The proposed method, Exploration-Boosted MC-PILCO (EB-MC-PILCO), leverages iLQR to generate informative, exploratory trajectories and initialize the policy, significantly reducing the number of required optimization steps. Experiments on the cart-pole task demonstrate that EB-MC-PILCO accelerates convergence compared to standard MC-PILCO, achieving up to $\bm{45.9\%}$ reduction in execution time when both methods solve the task in four trials. EB-MC-PILCO also maintains a $\bm{100\%}$ success rate across trials while solving the task faster, even in cases where MC-PILCO converges in fewer iterations.
title Accelerating Model-Based Reinforcement Learning using Non-Linear Trajectory Optimization
topic Machine Learning
Robotics
url https://arxiv.org/abs/2506.02767