VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiaoyu, Fu, Chaoyou, Yan, Chi, Wu, Chu, Gao, Haihan, Zhang, Yi-Fan, Dong, Shaoqi, Qian, Cheng, Luo, Bin, Yang, Xiuyong, Li, Guanwu, Cai, Yusheng, Shen, Yunhang, Jiang, Deqiang, Cao, Haoyu, Sun, Xing, Shan, Caifeng, He, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
by: Dong, Shaoqi, et al.
Published: (2025)
by: Dong, Shaoqi, et al.
Published: (2025)
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
by: Shen, Yunhang, et al.
Published: (2025)
by: Shen, Yunhang, et al.
Published: (2025)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
by: Fu, Chaoyou, et al.
Published: (2025)
by: Fu, Chaoyou, et al.
Published: (2025)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)
by: Li, Lijiang, et al.
Published: (2026)
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
by: Yang, Ruoliu, et al.
Published: (2026)
by: Yang, Ruoliu, et al.
Published: (2026)
VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
by: Long, Zuwei, et al.
Published: (2025)
by: Long, Zuwei, et al.
Published: (2025)
PersonaVLM: Long-Term Personalized Multimodal LLMs
by: Nie, Chang, et al.
Published: (2026)
by: Nie, Chang, et al.
Published: (2026)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
by: Fu, Chaoyou, et al.
Published: (2023)
by: Fu, Chaoyou, et al.
Published: (2023)
SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
"What I Sign Is Not What I See": Towards Explainable and Trustworthy Cryptocurrency Wallet Signatures
by: Qin, Yuyang, et al.
Published: (2026)
by: Qin, Yuyang, et al.
Published: (2026)
Asking Before Acting: Gather Information in Embodied Decision Making with Language Models
by: Chen, Xiaoyu, et al.
Published: (2023)
by: Chen, Xiaoyu, et al.
Published: (2023)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
by: Liu, Ruohan, et al.
Published: (2026)
by: Liu, Ruohan, et al.
Published: (2026)
HiLo: Learning Whole-Body Human-like Locomotion with Motion Tracking Controller
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
by: Fu, Chaoyou, et al.
Published: (2026)
by: Fu, Chaoyou, et al.
Published: (2026)
Pixelis: Reasoning in Pixels, from Seeing to Acting
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
by: Zhou, Chenyu, et al.
Published: (2024)
by: Zhou, Chenyu, et al.
Published: (2024)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
by: Wang, Xiong, et al.
Published: (2024)
by: Wang, Xiong, et al.
Published: (2024)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Partitioning triangle-free planar graphs into a forest and a linear forest
by: Liu, Guanwu, et al.
Published: (2025)
by: Liu, Guanwu, et al.
Published: (2025)
The Body Speaks: What Do Nurses Hear?
by: Lee SmithBattle, et al.
Published: (2025)
by: Lee SmithBattle, et al.
Published: (2025)
You Only Speak Once to See
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
by: Xing, Yazhou, et al.
Published: (2024)
by: Xing, Yazhou, et al.
Published: (2024)
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
by: Gao, Heting, et al.
Published: (2025)
by: Gao, Heting, et al.
Published: (2025)
See Me, Hear Me: Skype in the Classroom
by: Foote, Carolyn
Published: (2008)
by: Foote, Carolyn
Published: (2008)
VITA - Vocational Innovation through Teaching with AI
by: Ravotto, Pierfranco
Published: (2025)
by: Ravotto, Pierfranco
Published: (2025)
VITA: Vision-to-Action Flow Matching Policy
by: Gao, Dechen, et al.
Published: (2025)
by: Gao, Dechen, et al.
Published: (2025)
The general property of the tensor gravitational memory effect in theories of gravity
by: Hou, Shaoqi
Published: (2024)
by: Hou, Shaoqi
Published: (2024)
An improved lower bound for star-shaped Kakeya sets
by: Li, Shaoqi
Published: (2025)
by: Li, Shaoqi
Published: (2025)
Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
by: Wang, Shaoqi, et al.
Published: (2024)
by: Wang, Shaoqi, et al.
Published: (2024)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
by: Yin, Shukang, et al.
Published: (2024)
by: Yin, Shukang, et al.
Published: (2024)
See and Think: Embodied Agent in Virtual Environment
by: Zhao, Zhonghan, et al.
Published: (2023)
by: Zhao, Zhonghan, et al.
Published: (2023)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
by: Yin, Shukang, et al.
Published: (2023)
by: Yin, Shukang, et al.
Published: (2023)
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
by: Xu, Jiacheng, et al.
Published: (2026)
by: Xu, Jiacheng, et al.
Published: (2026)
On the Rise of Slurry Electrolysis for Energy Applications
by: Jingjing Xiong, et al.
Published: (2025)
by: Jingjing Xiong, et al.
Published: (2025)
Non-equilibrium dynamic hyperuniform states
by: Lei, Yusheng, et al.
Published: (2024)
by: Lei, Yusheng, et al.
Published: (2024)
How does a hyperuniform fluid freeze?
by: Lei, Yusheng, et al.
Published: (2023)
by: Lei, Yusheng, et al.
Published: (2023)
Metastable Hyperuniformity at Discontinuous Absorbing Transitions
by: Lei, Yusheng, et al.
Published: (2026)
by: Lei, Yusheng, et al.
Published: (2026)
Similar Items
-
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
by: Dong, Shaoqi, et al.
Published: (2025) -
VITA: Towards Open-Source Interactive Omni Multimodal LLM
by: Fu, Chaoyou, et al.
Published: (2024) -
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
by: Shen, Yunhang, et al.
Published: (2025) -
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
by: Fu, Chaoyou, et al.
Published: (2025) -
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
by: Li, Lijiang, et al.
Published: (2026)