Unified Video Action Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shuang, Gao, Yihuai, Sadigh, Dorsa, Song, Shuran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MotIF: Motion Instruction Fine-tuning
von: Hwang, Minyoung, et al.
Veröffentlicht: (2024)
von: Hwang, Minyoung, et al.
Veröffentlicht: (2024)
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
Geometry-aware 4D Video Generation for Robot Manipulation
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)
DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
von: Bauer, Dominik, et al.
Veröffentlicht: (2024)
von: Bauer, Dominik, et al.
Veröffentlicht: (2024)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
von: Liang, Junbang, et al.
Veröffentlicht: (2024)
What's the Move? Hybrid Imitation Learning via Salient Points
von: Sundaresan, Priya, et al.
Veröffentlicht: (2024)
von: Sundaresan, Priya, et al.
Veröffentlicht: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
von: Wu, Jimmy, et al.
Veröffentlicht: (2024)
von: Wu, Jimmy, et al.
Veröffentlicht: (2024)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
von: Guo, Jun, et al.
Veröffentlicht: (2026)
von: Guo, Jun, et al.
Veröffentlicht: (2026)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
von: Yan, Haodong, et al.
Veröffentlicht: (2026)
von: Yan, Haodong, et al.
Veröffentlicht: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
Differentiable Robot Rendering
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
von: Liu, Ruoshi, et al.
Veröffentlicht: (2024)
Unifying Language-Action Understanding and Generation for Autonomous Driving
von: Wang, Xinyang, et al.
Veröffentlicht: (2026)
von: Wang, Xinyang, et al.
Veröffentlicht: (2026)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
von: Zhao, Jianbo, et al.
Veröffentlicht: (2025)
von: Zhao, Jianbo, et al.
Veröffentlicht: (2025)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
von: Li, Peiyan, et al.
Veröffentlicht: (2026)
DriveVA: Video Action Models are Zero-Shot Drivers
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
von: Liu, Mengmeng, et al.
Veröffentlicht: (2026)
Precise Action-to-Video Generation Through Visual Action Prompts
von: Wang, Yuang, et al.
Veröffentlicht: (2025)
von: Wang, Yuang, et al.
Veröffentlicht: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Explore until Confident: Efficient Exploration for Embodied Question Answering
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
von: Ren, Allen Z., et al.
Veröffentlicht: (2024)
RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
Video Generation with Learned Action Prior
von: Sarkar, Meenakshi, et al.
Veröffentlicht: (2024)
von: Sarkar, Meenakshi, et al.
Veröffentlicht: (2024)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
ContactHandover: Contact-Guided Robot-to-Human Object Handover
von: Wang, Zixi, et al.
Veröffentlicht: (2024)
von: Wang, Zixi, et al.
Veröffentlicht: (2024)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
von: Ding, Hongyu, et al.
Veröffentlicht: (2026)
von: Ding, Hongyu, et al.
Veröffentlicht: (2026)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
von: Morin, Sacha, et al.
Veröffentlicht: (2025)
von: Morin, Sacha, et al.
Veröffentlicht: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
von: Lang, Xiaolei, et al.
Veröffentlicht: (2026)
von: Lang, Xiaolei, et al.
Veröffentlicht: (2026)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
Motus: A Unified Latent Action World Model
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
von: Yin, Zhenhan, et al.
Veröffentlicht: (2025)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MotIF: Motion Instruction Fine-tuning
von: Hwang, Minyoung, et al.
Veröffentlicht: (2024) -
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023) -
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025) -
Geometry-aware 4D Video Generation for Robot Manipulation
von: Liu, Zeyi, et al.
Veröffentlicht: (2025) -
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
von: Zhang, Jiazhao, et al.
Veröffentlicht: (2024)