Do World Action Models Generalize Better than VLAs? A Robustness Study
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhanguang, Li, Zhiyuan, Rahmati, Behnam, Yang, Rui Heng, Ma, Yintao, Rasouli, Amir, Pakdamansavoji, Sajjad, Wu, Yangzheng, Zhang, Lingfeng, Cao, Tongtong, Wen, Feng, Wang, Xinyu, Quan, Xingyue, Zhang, Yingxue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How VLAs (Really) Work In Open-World Environments
by: Rasouli, Amir, et al.
Published: (2026)
by: Rasouli, Amir, et al.
Published: (2026)
WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
by: Ma, Yintao, et al.
Published: (2025)
by: Ma, Yintao, et al.
Published: (2025)
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
by: Rasouli, Amir, et al.
Published: (2025)
by: Rasouli, Amir, et al.
Published: (2025)
Improving Robotic Manipulation Robustness via NICE Scene Surgery
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
by: Huang, Jinbang, et al.
Published: (2026)
by: Huang, Jinbang, et al.
Published: (2026)
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
by: Huang, Jinbang, et al.
Published: (2025)
by: Huang, Jinbang, et al.
Published: (2025)
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
by: Liu, Yuecheng, et al.
Published: (2025)
by: Liu, Yuecheng, et al.
Published: (2025)
ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
Impact Evaluation of Wastewater Treatment Based on the Anaerobic Digestion of Sewage Sludge Using the Life Cycle Assessment Method
by: Mohammad Rahmati, et al.
Published: (2024)
by: Mohammad Rahmati, et al.
Published: (2024)
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
by: Liu, Yuecheng, et al.
Published: (2025)
by: Liu, Yuecheng, et al.
Published: (2025)
Reinforcing VLAs in Task-Agnostic World Models
by: Wang, Yucen, et al.
Published: (2026)
by: Wang, Yucen, et al.
Published: (2026)
Pseudo-keypoint RKHS Learning for Self-supervised 6DoF Pose Estimation
by: Wu, Yangzheng, et al.
Published: (2023)
by: Wu, Yangzheng, et al.
Published: (2023)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
by: Cotnareanu, Joseph, et al.
Published: (2024)
by: Cotnareanu, Joseph, et al.
Published: (2024)
Diving Deeper Into Pedestrian Behavior Understanding: Intention Estimation, Action Prediction, and Event Risk Assessment
by: Rasouli, Amir, et al.
Published: (2024)
by: Rasouli, Amir, et al.
Published: (2024)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
How Do VLAs Effectively Inherit from VLMs?
by: Zhang, Chuheng, et al.
Published: (2025)
by: Zhang, Chuheng, et al.
Published: (2025)
E2ESlack: An End-to-End Graph-Based Framework for Pre-Routing Slack Prediction
by: Bodhe, Saurabh, et al.
Published: (2025)
by: Bodhe, Saurabh, et al.
Published: (2025)
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
by: Huang, Jinbang, et al.
Published: (2025)
by: Huang, Jinbang, et al.
Published: (2025)
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
Persuasion in Networks: Can the Sender Do Better than Using Public Signals?
by: Zhang, Yifan
Published: (2024)
by: Zhang, Yifan
Published: (2024)
TrACT: A Training Dynamics Aware Contrastive Learning Framework for Long-tail Trajectory Prediction
by: Zhang, Junrui, et al.
Published: (2024)
by: Zhang, Junrui, et al.
Published: (2024)
Ground Plane Projection for Improved Traffic Analytics at Intersections
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
by: Yuan, Haoran, et al.
Published: (2026)
by: Yuan, Haoran, et al.
Published: (2026)
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
by: Liu, Yushan, et al.
Published: (2026)
by: Liu, Yushan, et al.
Published: (2026)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
by: Zhang, Xinyu
Published: (2026)
by: Zhang, Xinyu
Published: (2026)
SnapFlow: One-Step Action Generation for Flow-Matching VLAs via Progressive Self-Distillation
by: Luan, Wuyang, et al.
Published: (2026)
by: Luan, Wuyang, et al.
Published: (2026)
Model Development and Simulation of Anaerobic Digestion Process Using Aspen Plus
by: Amir Sheikhi, et al.
Published: (2025)
by: Amir Sheikhi, et al.
Published: (2025)
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
by: Hancock, Asher J., et al.
Published: (2025)
by: Hancock, Asher J., et al.
Published: (2025)
When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
by: Aghaei, Matin, et al.
Published: (2025)
by: Aghaei, Matin, et al.
Published: (2025)
Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations
by: Kotseruba, Iuliia, et al.
Published: (2020)
by: Kotseruba, Iuliia, et al.
Published: (2020)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
by: Pai, Jonas, et al.
Published: (2025)
by: Pai, Jonas, et al.
Published: (2025)
Executable Code Actions Elicit Better LLM Agents
by: Wang, Xingyao, et al.
Published: (2024)
by: Wang, Xingyao, et al.
Published: (2024)
Byzantine-Robust Federated Learning Framework with Post-Quantum Secure Aggregation for Real-Time Threat Intelligence Sharing in Critical IoT Infrastructure
by: Rahmati, Milad, et al.
Published: (2026)
by: Rahmati, Milad, et al.
Published: (2026)
CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision Avoidance
by: Yang, Rui Heng, et al.
Published: (2025)
by: Yang, Rui Heng, et al.
Published: (2025)
Forgetting Similar Samples: Can Machine Unlearning Do it Better?
by: Xu, Heng, et al.
Published: (2026)
by: Xu, Heng, et al.
Published: (2026)
Taking a Bite Out of the Forbidden Fruit: Characterizing Third-Party Iranian iOS App Stores
by: Khanlari, Amirhossein, et al.
Published: (2026)
by: Khanlari, Amirhossein, et al.
Published: (2026)
Anansi: Scalable Characterization of Message-Based Job Scams
by: Pitumpe, Abisheka, et al.
Published: (2026)
by: Pitumpe, Abisheka, et al.
Published: (2026)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
by: Huang, Helong, et al.
Published: (2025)
by: Huang, Helong, et al.
Published: (2025)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
by: Moravej, Reza, et al.
Published: (2024)
by: Moravej, Reza, et al.
Published: (2024)
Similar Items
-
How VLAs (Really) Work In Open-World Environments
by: Rasouli, Amir, et al.
Published: (2026) -
WALDO: Where Unseen Model-based 6D Pose Estimation Meets Occlusion
by: Pakdamansavoji, Sajjad, et al.
Published: (2025) -
Box6D : Zero-shot Category-level 6D Pose Estimation of Warehouse Boxes
by: Ma, Yintao, et al.
Published: (2025) -
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
by: Rasouli, Amir, et al.
Published: (2025) -
Improving Robotic Manipulation Robustness via NICE Scene Surgery
by: Pakdamansavoji, Sajjad, et al.
Published: (2025)