VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Weiqi, Zhang, Quande, Zhai, Ruifeng, Lin, Liang, Wang, Guangrun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
by: Yin, Cheng, et al.
Published: (2025)
by: Yin, Cheng, et al.
Published: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
by: Shen, Junzhe, et al.
Published: (2026)
by: Shen, Junzhe, et al.
Published: (2026)
Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models
by: Jin, Ruofan, et al.
Published: (2026)
by: Jin, Ruofan, et al.
Published: (2026)
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
by: Tang, Yanchuan, et al.
Published: (2026)
by: Tang, Yanchuan, et al.
Published: (2026)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
by: Xiong, Zheng, et al.
Published: (2025)
by: Xiong, Zheng, et al.
Published: (2025)
CRL-VLA: Continual Vision-Language-Action Learning
by: Zeng, Qixin, et al.
Published: (2026)
by: Zeng, Qixin, et al.
Published: (2026)
On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability
by: Wang, Kevin, et al.
Published: (2024)
by: Wang, Kevin, et al.
Published: (2024)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
by: Yue, Yang, et al.
Published: (2024)
by: Yue, Yang, et al.
Published: (2024)
Curriculum Is More Influential Than Haptic Information During Reinforcement Learning of Object Manipulation Against Gravity
by: Ojaghi, Pegah, et al.
Published: (2024)
by: Ojaghi, Pegah, et al.
Published: (2024)
Watch Less, Feel More: Sim-to-Real RL for Generalizable Articulated Object Manipulation via Motion Adaptation and Impedance Control
by: Do, Tan-Dzung, et al.
Published: (2025)
by: Do, Tan-Dzung, et al.
Published: (2025)
Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study
by: Wei, Yubai, et al.
Published: (2026)
by: Wei, Yubai, et al.
Published: (2026)
VeTraSS: Vehicle Trajectory Similarity Search Through Graph Modeling and Representation Learning
by: Cheng, Ming, et al.
Published: (2024)
by: Cheng, Ming, et al.
Published: (2024)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling
by: Chen, Hongyu, et al.
Published: (2026)
by: Chen, Hongyu, et al.
Published: (2026)
RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models
by: Liufu, Weijia, et al.
Published: (2026)
by: Liufu, Weijia, et al.
Published: (2026)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
by: Bu, Qingwen, et al.
Published: (2025)
by: Bu, Qingwen, et al.
Published: (2025)
PhyPlan: Generalizable and Rapid Physical Task Planning with Physics Informed Skill Networks for Robot Manipulators
by: Chopra, Mudit, et al.
Published: (2024)
by: Chopra, Mudit, et al.
Published: (2024)
Distilling and Retrieving Generalizable Knowledge for Robot Manipulation via Language Corrections
by: Zha, Lihan, et al.
Published: (2023)
by: Zha, Lihan, et al.
Published: (2023)
ResNets Are Deeper Than You Think
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2025)
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
by: Zhang, Yixue, et al.
Published: (2026)
by: Zhang, Yixue, et al.
Published: (2026)
Conditioning Matters: Training Diffusion Policies is Faster Than You Think
by: Dong, Zibin, et al.
Published: (2025)
by: Dong, Zibin, et al.
Published: (2025)
Atomic Action Slicing: Planner-Aligned Options for Generalist VLA Agents
by: Tabakov, Stefan, et al.
Published: (2025)
by: Tabakov, Stefan, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
Revisiting Safe Exploration in Safe Reinforcement learning
by: Eckel, David, et al.
Published: (2024)
by: Eckel, David, et al.
Published: (2024)
Unicorn: A Universal and Collaborative Reinforcement Learning Approach Towards Generalizable Network-Wide Traffic Signal Control
by: Zhang, Yifeng, et al.
Published: (2025)
by: Zhang, Yifeng, et al.
Published: (2025)
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
by: Liu, Jun, et al.
Published: (2026)
by: Liu, Jun, et al.
Published: (2026)
Reasoning with Sampling: Your Base Model is Smarter Than You Think
by: Karan, Aayush, et al.
Published: (2025)
by: Karan, Aayush, et al.
Published: (2025)
Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models
by: Zhang, Yanyan, et al.
Published: (2026)
by: Zhang, Yanyan, et al.
Published: (2026)
Is Diversity All You Need for Scalable Robotic Manipulation?
by: Shi, Modi, et al.
Published: (2025)
by: Shi, Modi, et al.
Published: (2025)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters
by: Kong, Lingxiao, et al.
Published: (2026)
by: Kong, Lingxiao, et al.
Published: (2026)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Eurekaverse: Environment Curriculum Generation via Large Language Models
by: Liang, William, et al.
Published: (2024)
by: Liang, William, et al.
Published: (2024)
LoRA Is Slower Than You Think
by: Ko, Seokmin
Published: (2025)
by: Ko, Seokmin
Published: (2025)
Similar Items
-
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025) -
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
by: Yin, Cheng, et al.
Published: (2025) -
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
by: Shen, Junzhe, et al.
Published: (2026) -
Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models
by: Jin, Ruofan, et al.
Published: (2026) -
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
by: Tang, Yanchuan, et al.
Published: (2026)