Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Wenli, Lin, Haotian, Peng, Andy, Xue, Haoru, He, Tairan, Xie, Yuqi, Hu, Fengyuan, Wu, Jimmy, Luo, Zhengyi, Fan, Linxi "Jim", Shi, Guanya, Zhu, Yuke |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
by: He, Tairan, et al.
Published: (2025)
by: He, Tairan, et al.
Published: (2025)
HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots
by: He, Tairan, et al.
Published: (2024)
by: He, Tairan, et al.
Published: (2024)
Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation
by: He, Tairan, et al.
Published: (2024)
by: He, Tairan, et al.
Published: (2024)
WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)
by: Xiao, Wenli, et al.
Published: (2023)
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
by: Fu, Max, et al.
Published: (2026)
by: Fu, Max, et al.
Published: (2026)
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
AnyCar to Anywhere: Learning Universal Dynamics Model for Agile and Adaptive Mobility
by: Xiao, Wenli, et al.
Published: (2024)
by: Xiao, Wenli, et al.
Published: (2024)
Agile Mobility with Rapid Online Adaptation via Meta-learning and Uncertainty-aware MPPI
by: Kalaria, Dvij, et al.
Published: (2024)
by: Kalaria, Dvij, et al.
Published: (2024)
OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
by: He, Tairan, et al.
Published: (2024)
by: He, Tairan, et al.
Published: (2024)
ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
by: He, Tairan, et al.
Published: (2025)
by: He, Tairan, et al.
Published: (2025)
Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion
by: He, Tairan, et al.
Published: (2024)
by: He, Tairan, et al.
Published: (2024)
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing
by: Xue, Haoru, et al.
Published: (2024)
by: Xue, Haoru, et al.
Published: (2024)
HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos
by: Weng, Haoyang, et al.
Published: (2025)
by: Weng, Haoyang, et al.
Published: (2025)
Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics
by: Zhong, Yichao, et al.
Published: (2025)
by: Zhong, Yichao, et al.
Published: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
by: Grigsby, Jake, et al.
Published: (2023)
by: Grigsby, Jake, et al.
Published: (2023)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
by: Lin, Toru, et al.
Published: (2025)
by: Lin, Toru, et al.
Published: (2025)
TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint
by: Lin, Haotian, et al.
Published: (2025)
by: Lin, Haotian, et al.
Published: (2025)
Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control
by: Li, Yitang, et al.
Published: (2025)
by: Li, Yitang, et al.
Published: (2025)
Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning
by: Sobanbabu, Nikhil, et al.
Published: (2025)
by: Sobanbabu, Nikhil, et al.
Published: (2025)
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
by: Jiang, Zhenyu, et al.
Published: (2024)
by: Jiang, Zhenyu, et al.
Published: (2024)
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
RAMBO: RL-Augmented Model-Based Whole-Body Control for Loco-Manipulation
by: Cheng, Jin, et al.
Published: (2025)
by: Cheng, Jin, et al.
Published: (2025)
Self-Supervised Meta-Learning for All-Layer DNN-Based Adaptive Control with Stability Guarantees
by: He, Guanqi, et al.
Published: (2024)
by: He, Guanqi, et al.
Published: (2024)
ResWM: Residual-Action World Model for Visual RL
by: Zhang, Jseen, et al.
Published: (2026)
by: Zhang, Jseen, et al.
Published: (2026)
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
by: Lin, Kevin, et al.
Published: (2026)
by: Lin, Kevin, et al.
Published: (2026)
EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
by: Zheng, Ruijie, et al.
Published: (2026)
by: Zheng, Ruijie, et al.
Published: (2026)
NitroGen: An Open Foundation Model for Generalist Gaming Agents
by: Magne, Loïc, et al.
Published: (2026)
by: Magne, Loïc, et al.
Published: (2026)
World Action Models are Zero-shot Policies
by: Ye, Seonghyeon, et al.
Published: (2026)
by: Ye, Seonghyeon, et al.
Published: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
by: Maddukuri, Abhiram, et al.
Published: (2025)
by: Maddukuri, Abhiram, et al.
Published: (2025)
$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
Similar Items
-
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
by: Xue, Haoru, et al.
Published: (2025) -
VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
by: He, Tairan, et al.
Published: (2025) -
HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots
by: He, Tairan, et al.
Published: (2024) -
Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation
by: He, Tairan, et al.
Published: (2024) -
WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts
by: Zhang, Chong, et al.
Published: (2024)