Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yi, Li, Xinchen, Xie, Pengwei, Yang, Pu, Nie, Buqing, Cai, Yunuo, Zhang, Qinglin, Qu, Chendi, Wu, Jeffrey, Song, Jianheng, Ren, Xinlin, Huang, Jingshun, Pan, Mingjie, Feng, Siyuan, Chen, Zhi, Luo, Jianlan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
by: Pan, Mingjie, et al.
Published: (2026)
by: Pan, Mingjie, et al.
Published: (2026)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
by: Nie, Buqing, et al.
Published: (2025)
by: Nie, Buqing, et al.
Published: (2025)
Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization
by: Nie, Buqing, et al.
Published: (2025)
by: Nie, Buqing, et al.
Published: (2025)
Can Tabular Foundation Models Guide Exploration in Robot Policy Learning?
by: Ou, Buqing, et al.
Published: (2026)
by: Ou, Buqing, et al.
Published: (2026)
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
Where-to-Learn: Analytical Policy Gradient Directed Exploration for On-Policy Robotic Reinforcement Learning
by: Chang, Leixin, et al.
Published: (2026)
by: Chang, Leixin, et al.
Published: (2026)
Robot Fleet Learning via Policy Merging
by: Wang, Lirui, et al.
Published: (2023)
by: Wang, Lirui, et al.
Published: (2023)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Model Selection for Inverse Reinforcement Learning via Structural Risk Minimization
by: Qu, Chendi, et al.
Published: (2023)
by: Qu, Chendi, et al.
Published: (2023)
Disturbance-Aware Adaptive Compensation in Hybrid Force-Position Locomotion Policy for Legged Robots
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Effective Tuning Strategies for Generalist Robot Manipulation Policies
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
Learning Motion Skills with Adaptive Assistive Curriculum Force in Humanoid Robots
by: Cao, Zhanxiang, et al.
Published: (2025)
by: Cao, Zhanxiang, et al.
Published: (2025)
Minimizing Acoustic Noise: Enhancing Quiet Locomotion for Quadruped Robots in Indoor Applications
by: Cao, Zhanxiang, et al.
Published: (2025)
by: Cao, Zhanxiang, et al.
Published: (2025)
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025)
by: Xing, Youguang, et al.
Published: (2025)
RLIF: Interactive Imitation Learning as Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2023)
by: Luo, Jianlan, et al.
Published: (2023)
Select before Act: Spatially Decoupled Action Repetition for Continuous Control
by: Nie, Buqing, et al.
Published: (2025)
by: Nie, Buqing, et al.
Published: (2025)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2026)
by: Han, Xinchen, et al.
Published: (2026)
Energy‐Latency Tradeoffs for Service Placement Based on Reinforcement Learning in Edge Computing
by: Bing Tang, et al.
Published: (2025)
by: Bing Tang, et al.
Published: (2025)
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
Principled Data Augmentation for Learning to Solve Quadratic Programming Problems
by: Qian, Chendi, et al.
Published: (2025)
by: Qian, Chendi, et al.
Published: (2025)
PIQL: Projective Implicit Q-Learning with Support Constraint for Offline Reinforcement Learning
by: Han, Xinchen, et al.
Published: (2025)
by: Han, Xinchen, et al.
Published: (2025)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Robo-taxi Fleet Coordination at Scale via Reinforcement Learning
by: Tresca, Luigi, et al.
Published: (2025)
by: Tresca, Luigi, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms
by: Wang, Xinlin, et al.
Published: (2026)
by: Wang, Xinlin, et al.
Published: (2026)
Multi-Task Interactive Robot Fleet Learning with Visual World Models
by: Liu, Huihan, et al.
Published: (2024)
by: Liu, Huihan, et al.
Published: (2024)
OpenBot-Fleet: A System for Collective Learning with Real Robots
by: Müller, Matthias, et al.
Published: (2024)
by: Müller, Matthias, et al.
Published: (2024)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
by: Hu, Yucheng, et al.
Published: (2024)
by: Hu, Yucheng, et al.
Published: (2024)
Single Agent Robust Deep Reinforcement Learning for Bus Fleet Control
by: Zhang, Yifan
Published: (2025)
by: Zhang, Yifan
Published: (2025)
3DIOC: Direct Data-Driven Inverse Optimal Control for LTI Systems
by: Qu, Chendi, et al.
Published: (2024)
by: Qu, Chendi, et al.
Published: (2024)
Measuring Policy Distance for Multi-Agent Reinforcement Learning
by: Hu, Tianyi, et al.
Published: (2024)
by: Hu, Tianyi, et al.
Published: (2024)
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
by: Reuss, Moritz, et al.
Published: (2025)
by: Reuss, Moritz, et al.
Published: (2025)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
by: Atreya, Pranav, et al.
Published: (2025)
by: Atreya, Pranav, et al.
Published: (2025)
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
by: Song, Yunzhou, et al.
Published: (2026)
by: Song, Yunzhou, et al.
Published: (2026)
PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies
by: Jain, Arhan, et al.
Published: (2025)
by: Jain, Arhan, et al.
Published: (2025)
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
by: Reizinger, Patrik, et al.
Published: (2025)
by: Reizinger, Patrik, et al.
Published: (2025)
Similar Items
-
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
by: Pan, Mingjie, et al.
Published: (2026) -
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024) -
Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
by: Nie, Buqing, et al.
Published: (2025) -
Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization
by: Nie, Buqing, et al.
Published: (2025) -
Can Tabular Foundation Models Guide Exploration in Robot Policy Learning?
by: Ou, Buqing, et al.
Published: (2026)