Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yin, Patrick, Westenbroek, Tyler, Zhang, Zhengyu, Tran, Joshua, Dagnino, Ignacio, Shilamkar, Eeshani, Mbiziwo-Tiapo, Numfor, Bagaria, Simran, Liu, Xinlei, Mullins, Galen, Kolobov, Andrey, Gupta, Abhishek
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918424897126400
author Yin, Patrick
Westenbroek, Tyler
Zhang, Zhengyu
Tran, Joshua
Dagnino, Ignacio
Shilamkar, Eeshani
Mbiziwo-Tiapo, Numfor
Bagaria, Simran
Liu, Xinlei
Mullins, Galen
Kolobov, Andrey
Gupta, Abhishek
author_facet Yin, Patrick
Westenbroek, Tyler
Zhang, Zhengyu
Tran, Joshua
Dagnino, Ignacio
Shilamkar, Eeshani
Mbiziwo-Tiapo, Numfor
Bagaria, Simran
Liu, Xinlei
Mullins, Galen
Kolobov, Andrey
Gupta, Abhishek
contents Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engineering to design rewards, curricula, and demonstrations. Even with this engineering, they often fail on long-horizon, contact-rich manipulation tasks and do not meaningfully scale with compute, as performance quickly saturates when training revisits the same narrow regions of state space. We introduce OmniReset, a simple and scalable framework that enables on-policy reinforcement learning to robustly solve a broad class of dexterous manipulation tasks using a single reward function, fixed algorithm hyperparameters, no curricula, and no human demonstrations. Our key insight is that long-horizon exploration can be dramatically simplified by using simulator resets to systematically expose the RL algorithm to the diverse set of robot-object interactions which underlie dexterous manipulation. OmniReset programmatically generates such resets with minimal human input, converting additional compute directly into broader behavioral coverage and continued performance gains. We show that OmniReset gracefully scales to long-horizon dexterous manipulation tasks beyond the capabilities of existing approaches and is able to learn robust policies over significantly wider ranges of initial conditions than baselines. Finally, we distill OmniReset into visuomotor policies which display robust retrying behavior and substantially higher success rates than baselines when transferred to the real world zero-shot. Project webpage: https://weirdlabuw.github.io/omnireset/
format Preprint
id arxiv_https___arxiv_org_abs_2603_15789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
Yin, Patrick
Westenbroek, Tyler
Zhang, Zhengyu
Tran, Joshua
Dagnino, Ignacio
Shilamkar, Eeshani
Mbiziwo-Tiapo, Numfor
Bagaria, Simran
Liu, Xinlei
Mullins, Galen
Kolobov, Andrey
Gupta, Abhishek
Robotics
Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engineering to design rewards, curricula, and demonstrations. Even with this engineering, they often fail on long-horizon, contact-rich manipulation tasks and do not meaningfully scale with compute, as performance quickly saturates when training revisits the same narrow regions of state space. We introduce OmniReset, a simple and scalable framework that enables on-policy reinforcement learning to robustly solve a broad class of dexterous manipulation tasks using a single reward function, fixed algorithm hyperparameters, no curricula, and no human demonstrations. Our key insight is that long-horizon exploration can be dramatically simplified by using simulator resets to systematically expose the RL algorithm to the diverse set of robot-object interactions which underlie dexterous manipulation. OmniReset programmatically generates such resets with minimal human input, converting additional compute directly into broader behavioral coverage and continued performance gains. We show that OmniReset gracefully scales to long-horizon dexterous manipulation tasks beyond the capabilities of existing approaches and is able to learn robust policies over significantly wider ranges of initial conditions than baselines. Finally, we distill OmniReset into visuomotor policies which display robust retrying behavior and substantially higher success rates than baselines when transferred to the real world zero-shot. Project webpage: https://weirdlabuw.github.io/omnireset/
title Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
topic Robotics
url https://arxiv.org/abs/2603.15789