What Can RL Bring to VLA Generalization? An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jijia, Gao, Feng, Wei, Bingwen, Chen, Xinlei, Liao, Qingmin, Wu, Yi, Yu, Chao, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network
by: Liu, Jijia, et al.
Published: (2025)
by: Liu, Jijia, et al.
Published: (2025)
Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study
by: Wei, Yubai, et al.
Published: (2026)
by: Wei, Yubai, et al.
Published: (2026)
What Matters in Learning A Zero-Shot Sim-to-Real RL Policy for Quadrotor Control? A Comprehensive Study
by: Chen, Jiayu, et al.
Published: (2024)
by: Chen, Jiayu, et al.
Published: (2024)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
by: Wang, Songsheng, et al.
Published: (2025)
by: Wang, Songsheng, et al.
Published: (2025)
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
by: Liu, Jijia, et al.
Published: (2023)
by: Liu, Jijia, et al.
Published: (2023)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
Hysteresis-Aware Neural Network Modeling and Whole-Body Reinforcement Learning Control of Soft Robots
by: Chen, Zongyuan, et al.
Published: (2025)
by: Chen, Zongyuan, et al.
Published: (2025)
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
by: Ma, Chi, et al.
Published: (2024)
by: Ma, Chi, et al.
Published: (2024)
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
by: Su, Jianhai, et al.
Published: (2025)
by: Su, Jianhai, et al.
Published: (2025)
$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
Refined Policy Distillation: From VLA Generalists to RL Experts
by: Jülg, Tobias, et al.
Published: (2025)
by: Jülg, Tobias, et al.
Published: (2025)
An Improved Empirical Fisher Approximation for Natural Gradient Descent
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
by: Yang, Ningyuan, et al.
Published: (2025)
by: Yang, Ningyuan, et al.
Published: (2025)
OmniDrones: An Efficient and Flexible Platform for Reinforcement Learning in Drone Control
by: Xu, Botian, et al.
Published: (2023)
by: Xu, Botian, et al.
Published: (2023)
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
by: Zhang, Xingxuan, et al.
Published: (2025)
by: Zhang, Xingxuan, et al.
Published: (2025)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
by: Dai, Weinan, et al.
Published: (2026)
by: Dai, Weinan, et al.
Published: (2026)
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
A Benchmark Study of Deep-RL Methods for Maximum Coverage Problems over Graphs
by: Liang, Zhicheng, et al.
Published: (2024)
by: Liang, Zhicheng, et al.
Published: (2024)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
by: Zhang, Yixue, et al.
Published: (2026)
by: Zhang, Yixue, et al.
Published: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
by: Liang, Zhixuan, et al.
Published: (2025)
by: Liang, Zhixuan, et al.
Published: (2025)
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
by: Zhao, Lirui, et al.
Published: (2023)
by: Zhao, Lirui, et al.
Published: (2023)
VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
What Can You Do When You Have Zero Rewards During RL?
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
by: Bagaria, Vaidehi, et al.
Published: (2026)
by: Bagaria, Vaidehi, et al.
Published: (2026)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
What Does Flow Matching Bring To TD Learning?
by: Agrawalla, Bhavya, et al.
Published: (2026)
by: Agrawalla, Bhavya, et al.
Published: (2026)
Early Period of Training Impacts Adaptation for Out-of-Distribution Generalization: An Empirical Study
by: Liu, Chen Cecilia, et al.
Published: (2024)
by: Liu, Chen Cecilia, et al.
Published: (2024)
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
by: Jabbour, Jason, et al.
Published: (2025)
by: Jabbour, Jason, et al.
Published: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
by: Gao, Pengfei, et al.
Published: (2025)
by: Gao, Pengfei, et al.
Published: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Focus On What Matters: Separated Models For Visual-Based RL Generalization
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Labeling Case Similarity based on Co-Citation of Legal Articles in Judgment Documents with Empirical Dispute-Based Evaluation
by: Liu, Chao-Lin, et al.
Published: (2025)
by: Liu, Chao-Lin, et al.
Published: (2025)
Bringing Clustering to MLL: Weakly-Supervised Clustering for Partial Multi-Label Learning
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
by: Zhang, Zili, et al.
Published: (2026)
by: Zhang, Zili, et al.
Published: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Amplifier: Bringing Attention to Neglected Low-Energy Components in Time Series Forecasting
by: Fei, Jingru, et al.
Published: (2025)
by: Fei, Jingru, et al.
Published: (2025)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Similar Items
-
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network
by: Liu, Jijia, et al.
Published: (2025) -
Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study
by: Wei, Yubai, et al.
Published: (2026) -
What Matters in Learning A Zero-Shot Sim-to-Real RL Policy for Quadrotor Control? A Comprehensive Study
by: Chen, Jiayu, et al.
Published: (2024) -
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
by: Wang, Songsheng, et al.
Published: (2025) -
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination
by: Liu, Jijia, et al.
Published: (2023)