ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zhongyi, Zhu, Yichen, Zhu, Minjie, Wen, Junjie, Liu, Ning, Xu, Zhiyuan, Meng, Weibin, Cheng, Ran, Peng, Yaxin, Shen, Chaomin, Feng, Feifei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
von: Zhu, Minjie, et al.
Veröffentlicht: (2025)
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
Object-Centric Instruction Augmentation for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
von: Zhu, Minjie, et al.
Veröffentlicht: (2024)
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
WorldEval: World Model as Real-World Robot Policies Evaluator
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
von: Li, Yaxuan, et al.
Veröffentlicht: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
von: Wu, Kun, et al.
Veröffentlicht: (2024)
von: Wu, Kun, et al.
Veröffentlicht: (2024)
Efficient Feature Fusion for UAV Object Detection
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
von: Wang, Xudong, et al.
Veröffentlicht: (2025)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
von: Zhu, Yichen, et al.
Veröffentlicht: (2025)
von: Zhu, Yichen, et al.
Veröffentlicht: (2025)
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
von: Li, Meng, et al.
Veröffentlicht: (2025)
von: Li, Meng, et al.
Veröffentlicht: (2025)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
von: Wang, Qiuyue, et al.
Veröffentlicht: (2026)
von: Wang, Qiuyue, et al.
Veröffentlicht: (2026)
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Qiyuan, et al.
Veröffentlicht: (2025)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
von: Yu, Wenda, et al.
Veröffentlicht: (2026)
von: Yu, Wenda, et al.
Veröffentlicht: (2026)
Embodiment Transfer Learning for Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
von: Guo, Heyu, et al.
Veröffentlicht: (2025)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2025)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
von: Apanasevich, I., et al.
Veröffentlicht: (2026)
von: Apanasevich, I., et al.
Veröffentlicht: (2026)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025) -
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024) -
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
von: Zhu, Minjie, et al.
Veröffentlicht: (2025) -
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
von: Wen, Junjie, et al.
Veröffentlicht: (2024) -
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)