Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Rushuai, Feng, Zhiyuan, Zhang, Tianxiang, Wang, Kaixin, Zhang, Chuheng, Zhao, Li, Su, Xiu, Chen, Yi, Bian, Jiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
How Do VLAs Effectively Inherit from VLMs?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
von: Peng, Yuanfang, et al.
Veröffentlicht: (2026)
von: Peng, Yuanfang, et al.
Veröffentlicht: (2026)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
von: Wan, Jiansong, et al.
Veröffentlicht: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
von: Luo, Hao, et al.
Veröffentlicht: (2025)
von: Luo, Hao, et al.
Veröffentlicht: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
von: Xu, Charles, et al.
Veröffentlicht: (2026)
von: Xu, Charles, et al.
Veröffentlicht: (2026)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
von: Yang, Siyuan, et al.
Veröffentlicht: (2025)
von: Yang, Siyuan, et al.
Veröffentlicht: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
von: Jie, Haoxiang, et al.
Veröffentlicht: (2026)
von: Jie, Haoxiang, et al.
Veröffentlicht: (2026)
Trajectory First: A Curriculum for Discovering Diverse Policies
von: Braun, Cornelius V., et al.
Veröffentlicht: (2025)
von: Braun, Cornelius V., et al.
Veröffentlicht: (2025)
Unified Noise Steering for Efficient Human-Guided VLA Adaptation
von: Lu, Junjie, et al.
Veröffentlicht: (2026)
von: Lu, Junjie, et al.
Veröffentlicht: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation
von: Chen, Jiaming, et al.
Veröffentlicht: (2025)
von: Chen, Jiaming, et al.
Veröffentlicht: (2025)
Inject Once Survive Later: Backdooring Vision-Language-Action Models to Persist Through Downstream Fine-tuning
von: Zhou, Jianyi, et al.
Veröffentlicht: (2026)
von: Zhou, Jianyi, et al.
Veröffentlicht: (2026)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
von: Zhong, Zhide, et al.
Veröffentlicht: (2026)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
von: Lyu, Mingyang, et al.
Veröffentlicht: (2025)
von: Lyu, Mingyang, et al.
Veröffentlicht: (2025)
Dexbotic: Open-Source Vision-Language-Action Toolbox
von: Xie, Bin, et al.
Veröffentlicht: (2025)
von: Xie, Bin, et al.
Veröffentlicht: (2025)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
von: Xiao, Wenli, et al.
Veröffentlicht: (2025)
IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2024)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
SafeMove-RL: A Certifiable Reinforcement Learning Framework for Dynamic Motion Constraints in Trajectory Planning
von: Liu, Tengfei, et al.
Veröffentlicht: (2025)
von: Liu, Tengfei, et al.
Veröffentlicht: (2025)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
von: Fan, Shichao, et al.
Veröffentlicht: (2025)
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
von: Li, Meng, et al.
Veröffentlicht: (2025)
von: Li, Meng, et al.
Veröffentlicht: (2025)
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
von: Yu, En, et al.
Veröffentlicht: (2026)
von: Yu, En, et al.
Veröffentlicht: (2026)
A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios
von: Yang, Ruining, et al.
Veröffentlicht: (2025)
von: Yang, Ruining, et al.
Veröffentlicht: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
von: Zhang, Hongyin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyin, et al.
Veröffentlicht: (2025)
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
von: Fang, Irving, et al.
Veröffentlicht: (2025)
von: Fang, Irving, et al.
Veröffentlicht: (2025)
Automated Parking Trajectory Generation Using Deep Reinforcement Learning
von: Zhang, Zheyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zheyu, et al.
Veröffentlicht: (2025)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
von: Hu, Yucheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2025) -
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025) -
How Do VLAs Effectively Inherit from VLMs?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025) -
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2026) -
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
von: Peng, Yuanfang, et al.
Veröffentlicht: (2026)