VLA-0: Building State-of-the-Art VLAs with Zero Modification
Fuente:
arXiv
Saved in:
| Main Authors: | Goyal, Ankit, Hadfield, Hugo, Yang, Xuning, Blukis, Valts, Ramos, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
by: Yang, Xuning, et al.
Published: (2026)
by: Yang, Xuning, et al.
Published: (2026)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024)
by: Zook, Alex, et al.
Published: (2024)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
Learning to Plan & Schedule with Reinforcement-Learned Bimanual Robot Skills
by: Wan, Weikang, et al.
Published: (2025)
by: Wan, Weikang, et al.
Published: (2025)
How Do VLAs Effectively Inherit from VLMs?
by: Zhang, Chuheng, et al.
Published: (2025)
by: Zhang, Chuheng, et al.
Published: (2025)
How VLAs (Really) Work In Open-World Environments
by: Rasouli, Amir, et al.
Published: (2026)
by: Rasouli, Amir, et al.
Published: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
by: Tang, Jiaming, et al.
Published: (2025)
by: Tang, Jiaming, et al.
Published: (2025)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
by: Zhao, Jianchao, et al.
Published: (2026)
by: Zhao, Jianchao, et al.
Published: (2026)
Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs
by: Priyadershi, Abhinaw, et al.
Published: (2026)
by: Priyadershi, Abhinaw, et al.
Published: (2026)
DroneVLA: VLA based Aerial Manipulation
by: Mehboob, Fawad, et al.
Published: (2026)
by: Mehboob, Fawad, et al.
Published: (2026)
VLANeXt: Recipes for Building Strong VLA Models
by: Wu, Xiao-Ming, et al.
Published: (2026)
by: Wu, Xiao-Ming, et al.
Published: (2026)
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
by: Choi, Andrew, et al.
Published: (2026)
by: Choi, Andrew, et al.
Published: (2026)
RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour
by: Serpiva, Valerii, et al.
Published: (2025)
by: Serpiva, Valerii, et al.
Published: (2025)
Sci-VLA: Agentic VLA Inference Plugin for Long-Horizon Tasks in Scientific Experiments
by: Pang, Yiwen, et al.
Published: (2026)
by: Pang, Yiwen, et al.
Published: (2026)
End-to-End Dexterous Arm-Hand VLA Policies via Shared Autonomy: VR Teleoperation Augmented by Autonomous Hand VLA Policy for Efficient Data Collection
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
by: Takagi, Yusuke, et al.
Published: (2026)
by: Takagi, Yusuke, et al.
Published: (2026)
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
by: Zhou, Hui, et al.
Published: (2025)
by: Zhou, Hui, et al.
Published: (2025)
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
by: Yan, Hongyu, et al.
Published: (2026)
by: Yan, Hongyu, et al.
Published: (2026)
MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
by: Zhong, Yiming, et al.
Published: (2026)
by: Zhong, Yiming, et al.
Published: (2026)
ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
by: Zhang, Zhuoyang, et al.
Published: (2026)
by: Zhang, Zhuoyang, et al.
Published: (2026)
Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion
by: Li, Zhuo, et al.
Published: (2025)
by: Li, Zhuo, et al.
Published: (2025)
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
by: Li, Haoyun, et al.
Published: (2025)
by: Li, Haoyun, et al.
Published: (2025)
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
by: Gubernatorov, Konstantin, et al.
Published: (2025)
by: Gubernatorov, Konstantin, et al.
Published: (2025)
WorldVLA: Towards Autoregressive Action World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
Beyond Task Success: Behavioral and Representational Diagnostics for WAM and VLA
by: Mai, Hung, et al.
Published: (2026)
by: Mai, Hung, et al.
Published: (2026)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
by: Zhang, Borong, et al.
Published: (2025)
by: Zhang, Borong, et al.
Published: (2025)
TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation
by: Yang, Kai, et al.
Published: (2026)
by: Yang, Kai, et al.
Published: (2026)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
by: Du, Zhiying, et al.
Published: (2025)
by: Du, Zhiying, et al.
Published: (2025)
VLA-RAIL: A Real-Time Asynchronous Inference Linker for VLA Models and Robots
by: Zhao, Yongsheng, et al.
Published: (2025)
by: Zhao, Yongsheng, et al.
Published: (2025)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
by: Zhang, Dapeng, et al.
Published: (2025)
by: Zhang, Dapeng, et al.
Published: (2025)
Similar Items
-
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025) -
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
by: Yang, Xuning, et al.
Published: (2026) -
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024) -
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024) -
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)