How Do VLAs Effectively Inherit from VLMs?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chuheng, Yang, Rushuai, Chen, Xiaoyu, Wang, Kaixin, Zhao, Li, Chen, Yi, Bian, Jiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2024)
How VLAs (Really) Work In Open-World Environments
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
What Do Latent Action Models Actually Learn?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
von: Yang, Rushuai, et al.
Veröffentlicht: (2025)
ARO: Large Language Model Supervised Robotics Text2Skill Autonomous Learning
von: Chen, Yiwen, et al.
Veröffentlicht: (2024)
von: Chen, Yiwen, et al.
Veröffentlicht: (2024)
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
von: Tang, Jiaming, et al.
Veröffentlicht: (2025)
Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
Lamarckian Inheritance in Dynamic Environments: How Key Variables Affect Evolutionary Dynamics
von: de Bruin, K. Ege, et al.
Veröffentlicht: (2026)
von: de Bruin, K. Ege, et al.
Veröffentlicht: (2026)
Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
von: Jian, Yingzhao, et al.
Veröffentlicht: (2025)
von: Jian, Yingzhao, et al.
Veröffentlicht: (2025)
A Collision-Aware Cable Grasping Method in Cluttered Environment
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
Reinforcing VLAs in Task-Agnostic World Models
von: Wang, Yucen, et al.
Veröffentlicht: (2026)
von: Wang, Yucen, et al.
Veröffentlicht: (2026)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
Neural Interaction Energy for Multi-Agent Trajectory Prediction
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
Lamarckian Inheritance Improves Robot Evolution in Dynamic Environments
von: Luo, Jie, et al.
Veröffentlicht: (2024)
von: Luo, Jie, et al.
Veröffentlicht: (2024)
Orchestrating Spatial Semantics via a Zone-Graph Paradigm for Intricate Indoor Scene Generation
von: Zhang, Meisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Meisheng, et al.
Veröffentlicht: (2026)
ContactDexNet: Multi-fingered Robotic Hand Grasping in Cluttered Environments through Hand-object Contact Semantic Mapping
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
von: Zhang, Lei, et al.
Veröffentlicht: (2024)
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
von: Kang, Sinjae, et al.
Veröffentlicht: (2026)
von: Kang, Sinjae, et al.
Veröffentlicht: (2026)
An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
von: Ji, Kailun, et al.
Veröffentlicht: (2025)
von: Ji, Kailun, et al.
Veröffentlicht: (2025)
Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
von: Chiang, Hao-Tien Lewis, et al.
Veröffentlicht: (2024)
von: Chiang, Hao-Tien Lewis, et al.
Veröffentlicht: (2024)
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
von: Yang, Fengze, et al.
Veröffentlicht: (2025)
von: Yang, Fengze, et al.
Veröffentlicht: (2025)
Do You Need Proprioceptive States in Visuomotor Policies?
von: Zhao, Juntu, et al.
Veröffentlicht: (2025)
von: Zhao, Juntu, et al.
Veröffentlicht: (2025)
Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences
von: Wu, Wenxi, et al.
Veröffentlicht: (2026)
von: Wu, Wenxi, et al.
Veröffentlicht: (2026)
PhysORD: A Neuro-Symbolic Approach for Physics-infused Motion Prediction in Off-road Driving
von: Zhao, Zhipeng, et al.
Veröffentlicht: (2024)
von: Zhao, Zhipeng, et al.
Veröffentlicht: (2024)
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
von: Guo, Yanjiang, et al.
Veröffentlicht: (2023)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2023)
Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
Prediction with Action: Visual Policy Learning via Joint Denoising Process
von: Guo, Yanjiang, et al.
Veröffentlicht: (2024)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2024)
Revealing Interpretable Failure Modes of VLMs
von: Chaudhary, Isha, et al.
Veröffentlicht: (2026)
von: Chaudhary, Isha, et al.
Veröffentlicht: (2026)
An LSTM Feature Imitation Network for Hand Movement Recognition from sEMG Signals
von: Wu, Chuheng, et al.
Veröffentlicht: (2024)
von: Wu, Chuheng, et al.
Veröffentlicht: (2024)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
PeriGuru: A Peripheral Robotic Mobile App Operation Assistant based on GUI Image Understanding and Prompting with LLM
von: Fu, Kelin, et al.
Veröffentlicht: (2024)
von: Fu, Kelin, et al.
Veröffentlicht: (2024)
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
von: Zhang, Jianke, et al.
Veröffentlicht: (2025)
von: Zhang, Jianke, et al.
Veröffentlicht: (2025)
Efficient and Generalized end-to-end Autonomous Driving System with Latent Deep Reinforcement Learning and Demonstrations
von: Tang, Zuojin, et al.
Veröffentlicht: (2024)
von: Tang, Zuojin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
von: Yang, Rushuai, et al.
Veröffentlicht: (2025) -
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025) -
IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2024) -
How VLAs (Really) Work In Open-World Environments
von: Rasouli, Amir, et al.
Veröffentlicht: (2026) -
What Do Latent Action Models Actually Learn?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)