Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hongyin, Zhang, Shiyuan, Jin, Junxi, Zeng, Qixin, Qiao, Yifan, Lu, Hongchao, Wang, Donglin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026)
by: Li, Runze, et al.
Published: (2026)
A Real-World Quadrupedal Locomotion Benchmark for Offline Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2023)
by: Zhang, Hongyin, et al.
Published: (2023)
CMR: Contractive Mapping Embeddings for Robust Humanoid Locomotion on Unstructured Terrains
by: Zeng, Qixin, et al.
Published: (2026)
by: Zeng, Qixin, et al.
Published: (2026)
CRL-VLA: Continual Vision-Language-Action Learning
by: Zeng, Qixin, et al.
Published: (2026)
by: Zeng, Qixin, et al.
Published: (2026)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
by: Xiao, Wei, et al.
Published: (2025)
by: Xiao, Wei, et al.
Published: (2025)
ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation
by: Zhang, Songyuan, et al.
Published: (2026)
by: Zhang, Songyuan, et al.
Published: (2026)
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
by: Jiang, Zhennan, et al.
Published: (2026)
by: Jiang, Zhennan, et al.
Published: (2026)
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning
by: Lee, Sungyoung, et al.
Published: (2026)
by: Lee, Sungyoung, et al.
Published: (2026)
Refined Policy Distillation: From VLA Generalists to RL Experts
by: Jülg, Tobias, et al.
Published: (2025)
by: Jülg, Tobias, et al.
Published: (2025)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
by: Zhong, Zhide, et al.
Published: (2026)
by: Zhong, Zhide, et al.
Published: (2026)
Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models
by: Jin, Ruofan, et al.
Published: (2026)
by: Jin, Ruofan, et al.
Published: (2026)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
by: Huang, Renming, et al.
Published: (2024)
by: Huang, Renming, et al.
Published: (2024)
LongNav-R1: Horizon-Adaptive Multi-Turn RL for Long-Horizon VLA Navigation
by: Hu, Yue, et al.
Published: (2026)
by: Hu, Yue, et al.
Published: (2026)
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
by: Bagaria, Vaidehi, et al.
Published: (2026)
by: Bagaria, Vaidehi, et al.
Published: (2026)
VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
by: Wang, Si-Cheng, et al.
Published: (2025)
by: Wang, Si-Cheng, et al.
Published: (2025)
DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Language-Conditioned Offline RL for Multi-Robot Navigation
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
by: Zhuang, Zifeng, et al.
Published: (2025)
by: Zhuang, Zifeng, et al.
Published: (2025)
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
by: Yue, Junpeng, et al.
Published: (2025)
by: Yue, Junpeng, et al.
Published: (2025)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
by: Li, Mingxuan, et al.
Published: (2026)
by: Li, Mingxuan, et al.
Published: (2026)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
by: Park, Seohong, et al.
Published: (2023)
by: Park, Seohong, et al.
Published: (2023)
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
by: Kim, Sungha, et al.
Published: (2026)
by: Kim, Sungha, et al.
Published: (2026)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
by: Xiao, Junjin, et al.
Published: (2025)
by: Xiao, Junjin, et al.
Published: (2025)
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
by: Su, Huikang, et al.
Published: (2025)
by: Su, Huikang, et al.
Published: (2025)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
by: Kim, Changyeon, et al.
Published: (2025)
by: Kim, Changyeon, et al.
Published: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Adaptive Q-Chunking for Offline-to-Online Reinforcement Learning
by: Gireesh, Nandiraju, et al.
Published: (2026)
by: Gireesh, Nandiraju, et al.
Published: (2026)
Learning from Mistakes: Post-Training for Driving VLA with Takeover Data
by: Gao, Yinfeng, et al.
Published: (2026)
by: Gao, Yinfeng, et al.
Published: (2026)
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
by: Peng, Yuanfang, et al.
Published: (2026)
by: Peng, Yuanfang, et al.
Published: (2026)
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
by: Rowe, Luke, et al.
Published: (2024)
by: Rowe, Luke, et al.
Published: (2024)
OpenVLA: An Open-Source Vision-Language-Action Model
by: Kim, Moo Jin, et al.
Published: (2024)
by: Kim, Moo Jin, et al.
Published: (2024)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
Similar Items
-
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025) -
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026) -
A Real-World Quadrupedal Locomotion Benchmark for Offline Reinforcement Learning
by: Zhang, Hongyin, et al.
Published: (2023) -
CMR: Contractive Mapping Embeddings for Robust Humanoid Locomotion on Unstructured Terrains
by: Zeng, Qixin, et al.
Published: (2026) -
CRL-VLA: Continual Vision-Language-Action Learning
by: Zeng, Qixin, et al.
Published: (2026)