Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cai, Shizhe, Yin, Zeya, Jacob, Jayadeep, Ramos, Fabio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908863825969152
author Cai, Shizhe
Yin, Zeya
Jacob, Jayadeep
Ramos, Fabio
author_facet Cai, Shizhe
Yin, Zeya
Jacob, Jayadeep
Ramos, Fabio
contents Model Predictive Control (MPC) enables reliable trajectory optimization under dynamics constraints, but often depends on accurate dynamics models and carefully hand-designed cost functions. Recent learning-based MPC methods aim to reduce these modeling and cost-design burdens by learning dynamics, priors, or value-related guidance signals. Yet many existing approaches still rely on deterministic gradient-based solvers (e.g., differentiable MPC) or parametric sampling-based updates (e.g., CEM/MPPI), which can lead to mode collapse and convergence to a single dominant solution. We propose Q-SVMPC, a Q-guided Stein variational MPC method with an RL-informed policy prior, which casts learning-based MPC as trajectory-level posterior inference and refines trajectory particles via SVGD under learned soft Q-value guidance to explicitly preserve diverse solutions. Experiments on navigation, robotic manipulation, and a real-world fruit-picking task show improved sample efficiency, stability, and robustness over MPC, model-free RL, and learning-based MPC baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
Cai, Shizhe
Yin, Zeya
Jacob, Jayadeep
Ramos, Fabio
Robotics
Artificial Intelligence
Machine Learning
Model Predictive Control (MPC) enables reliable trajectory optimization under dynamics constraints, but often depends on accurate dynamics models and carefully hand-designed cost functions. Recent learning-based MPC methods aim to reduce these modeling and cost-design burdens by learning dynamics, priors, or value-related guidance signals. Yet many existing approaches still rely on deterministic gradient-based solvers (e.g., differentiable MPC) or parametric sampling-based updates (e.g., CEM/MPPI), which can lead to mode collapse and convergence to a single dominant solution. We propose Q-SVMPC, a Q-guided Stein variational MPC method with an RL-informed policy prior, which casts learning-based MPC as trajectory-level posterior inference and refines trajectory particles via SVGD under learned soft Q-value guidance to explicitly preserve diverse solutions. Experiments on navigation, robotic manipulation, and a real-world fruit-picking task show improved sample efficiency, stability, and robustness over MPC, model-free RL, and learning-based MPC baselines.
title Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.06625