VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Wei, Ding, Pengxiang, Zhang, Min, Gong, Zhefei, Bai, Shuanghao, Zhao, Han, Wang, Donglin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
di: Fan, Yiguo, et al.
Pubblicazione: (2025)
di: Fan, Yiguo, et al.
Pubblicazione: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
di: Gong, Zhefei, et al.
Pubblicazione: (2025)
di: Gong, Zhefei, et al.
Pubblicazione: (2025)
ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
di: Zhao, Han, et al.
Pubblicazione: (2025)
di: Zhao, Han, et al.
Pubblicazione: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
Learning Robotic Policy with Imagined Transition: Mitigating the Trade-off between Robustness and Optimality
di: Xiao, Wei, et al.
Pubblicazione: (2025)
di: Xiao, Wei, et al.
Pubblicazione: (2025)
OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
di: Cui, Can, et al.
Pubblicazione: (2025)
di: Cui, Can, et al.
Pubblicazione: (2025)
VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
di: Zhang, Haoran, et al.
Pubblicazione: (2025)
di: Zhang, Haoran, et al.
Pubblicazione: (2025)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
di: Chen, Jiayi, et al.
Pubblicazione: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
di: Gong, Zhefei, et al.
Pubblicazione: (2024)
di: Gong, Zhefei, et al.
Pubblicazione: (2024)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
di: Wang, Yihao, et al.
Pubblicazione: (2025)
di: Wang, Yihao, et al.
Pubblicazione: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
di: Zhang, Hongyin, et al.
Pubblicazione: (2025)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
di: Li, Fuhao, et al.
Pubblicazione: (2025)
di: Li, Fuhao, et al.
Pubblicazione: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
di: Song, Wenxuan, et al.
Pubblicazione: (2024)
di: Song, Wenxuan, et al.
Pubblicazione: (2024)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
di: Lin, Minghui, et al.
Pubblicazione: (2025)
di: Lin, Minghui, et al.
Pubblicazione: (2025)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
di: Huang, Helong, et al.
Pubblicazione: (2025)
di: Huang, Helong, et al.
Pubblicazione: (2025)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
di: Zhao, Wentao, et al.
Pubblicazione: (2024)
di: Zhao, Wentao, et al.
Pubblicazione: (2024)
CRL-VLA: Continual Vision-Language-Action Learning
di: Zeng, Qixin, et al.
Pubblicazione: (2026)
di: Zeng, Qixin, et al.
Pubblicazione: (2026)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
di: Wei, Xiangyi, et al.
Pubblicazione: (2025)
Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
di: Fan, Cunxin, et al.
Pubblicazione: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
di: Yang, Shuai, et al.
Pubblicazione: (2025)
di: Yang, Shuai, et al.
Pubblicazione: (2025)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
di: Zhang, Yihao, et al.
Pubblicazione: (2025)
di: Zhang, Yihao, et al.
Pubblicazione: (2025)
Survey of Vision-Language-Action Models for Embodied Manipulation
di: Li, Haoran, et al.
Pubblicazione: (2025)
di: Li, Haoran, et al.
Pubblicazione: (2025)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
di: Zuo, Kuangji, et al.
Pubblicazione: (2026)
di: Zuo, Kuangji, et al.
Pubblicazione: (2026)
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
di: He, Yayun, et al.
Pubblicazione: (2026)
di: He, Yayun, et al.
Pubblicazione: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
di: Peng, Xiongfeng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
di: Zhao, Wei, et al.
Pubblicazione: (2025) -
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2025) -
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
di: Zhao, Han, et al.
Pubblicazione: (2025) -
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
di: Fan, Yiguo, et al.
Pubblicazione: (2025) -
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)