Constrained Sampling to Guide Universal Manipulation RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toussaint, Marc, Braun, Cornelius V., Cobo-Briesewitz, Eckart, Auddy, Sayantan, Jordana, Armand, Carpentier, Justin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910016528711680
author Toussaint, Marc
Braun, Cornelius V.
Cobo-Briesewitz, Eckart
Auddy, Sayantan
Jordana, Armand
Carpentier, Justin
author_facet Toussaint, Marc
Braun, Cornelius V.
Cobo-Briesewitz, Eckart
Auddy, Sayantan
Jordana, Armand
Carpentier, Justin
contents We consider how model-based solvers can be leveraged to guide training of a universal policy to control from any feasible start state to any feasible goal in a contact-rich manipulation setting. While Reinforcement Learning (RL) has demonstrated its strength in such settings, it may struggle to sufficiently explore and discover complex manipulation strategies, especially in sparse-reward settings. Our approach is based on the idea of a lower-dimensional manifold of feasible, likely-visited states during such manipulation and to guide RL with a sampler from this manifold. We propose Sample-Guided RL, which uses model-based constraint solvers to efficiently sample feasible configurations (satisfying differentiable collision, contact, and force constraints) and leverage them to guide RL for universal (goal-conditioned) manipulation policies. We study using this data directly to bias state visitation, as well as using black-box optimization of open-loop trajectories between random configurations to impose a state bias and optionally add a behavior cloning loss. In a minimalistic double sphere manipulation setting, Sample-Guided RL discovers complex manipulation strategies and achieves high success rates in reaching any statically stable state. In a more challenging panda arm setting, our approach achieves a significant success rate over a near-zero baseline, and demonstrates a breadth of complex whole-body-contact manipulation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08557
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Constrained Sampling to Guide Universal Manipulation RL
Toussaint, Marc
Braun, Cornelius V.
Cobo-Briesewitz, Eckart
Auddy, Sayantan
Jordana, Armand
Carpentier, Justin
Robotics
We consider how model-based solvers can be leveraged to guide training of a universal policy to control from any feasible start state to any feasible goal in a contact-rich manipulation setting. While Reinforcement Learning (RL) has demonstrated its strength in such settings, it may struggle to sufficiently explore and discover complex manipulation strategies, especially in sparse-reward settings. Our approach is based on the idea of a lower-dimensional manifold of feasible, likely-visited states during such manipulation and to guide RL with a sampler from this manifold. We propose Sample-Guided RL, which uses model-based constraint solvers to efficiently sample feasible configurations (satisfying differentiable collision, contact, and force constraints) and leverage them to guide RL for universal (goal-conditioned) manipulation policies. We study using this data directly to bias state visitation, as well as using black-box optimization of open-loop trajectories between random configurations to impose a state bias and optionally add a behavior cloning loss. In a minimalistic double sphere manipulation setting, Sample-Guided RL discovers complex manipulation strategies and achieves high success rates in reaching any statically stable state. In a more challenging panda arm setting, our approach achieves a significant success rate over a near-zero baseline, and demonstrates a breadth of complex whole-body-contact manipulation strategies.
title Constrained Sampling to Guide Universal Manipulation RL
topic Robotics
url https://arxiv.org/abs/2602.08557