Refined Policy Distillation: From VLA Generalists to RL Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Jülg, Tobias, Burgard, Wolfram, Walter, Florian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
by: Krack, Pierre, et al.
Published: (2026)
by: Krack, Pierre, et al.
Published: (2026)
Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness
by: Mirjalili, Reihaneh, et al.
Published: (2025)
by: Mirjalili, Reihaneh, et al.
Published: (2025)
VLAgents: A Policy Server for Efficient VLA Inference
by: Jülg, Tobias, et al.
Published: (2026)
by: Jülg, Tobias, et al.
Published: (2026)
FlowTouch: View-Invariant Visuo-Tactile Prediction
by: Bien, Seongjin, et al.
Published: (2026)
by: Bien, Seongjin, et al.
Published: (2026)
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
by: Jülg, Tobias, et al.
Published: (2025)
by: Jülg, Tobias, et al.
Published: (2025)
LLM-Pack: Intuitive Grocery Handling for Logistics Applications
by: Blei, Yannik, et al.
Published: (2025)
by: Blei, Yannik, et al.
Published: (2025)
Agent-Agnostic Centralized Training for Decentralized Multi-Agent Cooperative Driving
by: Yan, Shengchao, et al.
Published: (2024)
by: Yan, Shengchao, et al.
Published: (2024)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
VLM-Vac: Enhancing Smart Vacuums through VLM Knowledge Distillation and Language-Guided Experience Replay
by: Mirjalili, Reihaneh, et al.
Published: (2024)
by: Mirjalili, Reihaneh, et al.
Published: (2024)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Atomic Action Slicing: Planner-Aligned Options for Generalist VLA Agents
by: Tabakov, Stefan, et al.
Published: (2025)
by: Tabakov, Stefan, et al.
Published: (2025)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
DiWA: Diffusion Policy Adaptation with World Models
by: Chandra, Akshay L, et al.
Published: (2025)
by: Chandra, Akshay L, et al.
Published: (2025)
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
by: Huang, Chenguang, et al.
Published: (2024)
by: Huang, Chenguang, et al.
Published: (2024)
Effective Tuning Strategies for Generalist Robot Manipulation Policies
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
by: Bagaria, Vaidehi, et al.
Published: (2026)
by: Bagaria, Vaidehi, et al.
Published: (2026)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
by: Atreya, Pranav, et al.
Published: (2025)
by: Atreya, Pranav, et al.
Published: (2025)
PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies
by: Jain, Arhan, et al.
Published: (2025)
by: Jain, Arhan, et al.
Published: (2025)
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
by: Song, Yunzhou, et al.
Published: (2026)
by: Song, Yunzhou, et al.
Published: (2026)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies
by: Singh, Anya, et al.
Published: (2026)
by: Singh, Anya, et al.
Published: (2026)
$π^{*}_{0.6}$: a VLA That Learns From Experience
by: Intelligence, Physical, et al.
Published: (2025)
by: Intelligence, Physical, et al.
Published: (2025)
Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control
by: Peng, Quanquan, et al.
Published: (2026)
by: Peng, Quanquan, et al.
Published: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
End-to-end RL Improves Dexterous Grasping Policies
by: Singh, Ritvik, et al.
Published: (2025)
by: Singh, Ritvik, et al.
Published: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
by: Hirose, Noriaki, et al.
Published: (2026)
by: Hirose, Noriaki, et al.
Published: (2026)
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
by: Kim, Sungha, et al.
Published: (2026)
by: Kim, Sungha, et al.
Published: (2026)
One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation
by: Wang, Zhendong, et al.
Published: (2024)
by: Wang, Zhendong, et al.
Published: (2024)
Push Smarter, Not Harder: Hierarchical RL-Diffusion Policy for Efficient Nonprehensile Manipulation
by: Caro, Steven, et al.
Published: (2025)
by: Caro, Steven, et al.
Published: (2025)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
by: Huang, Chenguang, et al.
Published: (2025)
by: Huang, Chenguang, et al.
Published: (2025)
LUMOS: Language-Conditioned Imitation Learning with World Models
by: Nematollahi, Iman, et al.
Published: (2025)
by: Nematollahi, Iman, et al.
Published: (2025)
Fisher Decorator: Refining Flow Policy via a Local Transport Map
by: Cheng, Xiaoyuan, et al.
Published: (2026)
by: Cheng, Xiaoyuan, et al.
Published: (2026)
Similar Items
-
Rewarding DINO: Predicting Dense Rewards with Vision Foundation Models
by: Krack, Pierre, et al.
Published: (2026) -
Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness
by: Mirjalili, Reihaneh, et al.
Published: (2025) -
VLAgents: A Policy Server for Efficient VLA Inference
by: Jülg, Tobias, et al.
Published: (2026) -
FlowTouch: View-Invariant Visuo-Tactile Prediction
by: Bien, Seongjin, et al.
Published: (2026) -
Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
by: Jülg, Tobias, et al.
Published: (2025)