Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Siyin, Yu, Wenyi, Chen, Xianzhao, Tian, Xiaohai, Zhang, Jun, Lu, Lu, Zhang, Chao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2510.16756
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908979077054464
author Wang, Siyin
Yu, Wenyi
Chen, Xianzhao
Tian, Xiaohai
Zhang, Jun
Lu, Lu
Zhang, Chao
author_facet Wang, Siyin
Yu, Wenyi
Chen, Xianzhao
Tian, Xiaohai
Zhang, Jun
Lu, Lu
Zhang, Chao
contents Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabilities is essential for building models simulating humans. We present ELLSA (End-to-end Listen, Look, Speak and Act), which, to our knowledge, is the first full-duplex, end-to-end model that simultaneously perceives and generates across vision, text, speech, and action within a single architecture, enabling interaction patterns previously out of reach, yielding more natural, human-like behaviors. At its core is a novel SA-MoE architecture (Self-Attention Mixture-of-Experts) that routes each modality to specialized experts and fuses them through a unified attention backbone. This provides a generalizable solution for joint multimodal perception and concurrent generation, leveraging strong pre-trained components while enabling efficient modality integration and mitigating modality interference. On speech-interaction and robot-manipulation benchmarks, ELLSA matches modality-specific baselines, while uniquely supporting advanced multimodal and full-duplex behaviors such as dialogue and action turn-taking, defective instruction rejection, speaking-while-acting, context-grounded visual question answering, and action barge-ins. We contend that ELLSA represents a step toward more natural and general interactive intelligence, contributing to the broader pursuit of artificial general intelligence. All data, code and model checkpoints will be released at https://github.com/bytedance/SALMONN/tree/ELLSA.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle End-to-end Listen, Look, Speak and Act
Wang, Siyin
Yu, Wenyi
Chen, Xianzhao
Tian, Xiaohai
Zhang, Jun
Lu, Lu
Zhang, Chao
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
Audio and Speech Processing
Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabilities is essential for building models simulating humans. We present ELLSA (End-to-end Listen, Look, Speak and Act), which, to our knowledge, is the first full-duplex, end-to-end model that simultaneously perceives and generates across vision, text, speech, and action within a single architecture, enabling interaction patterns previously out of reach, yielding more natural, human-like behaviors. At its core is a novel SA-MoE architecture (Self-Attention Mixture-of-Experts) that routes each modality to specialized experts and fuses them through a unified attention backbone. This provides a generalizable solution for joint multimodal perception and concurrent generation, leveraging strong pre-trained components while enabling efficient modality integration and mitigating modality interference. On speech-interaction and robot-manipulation benchmarks, ELLSA matches modality-specific baselines, while uniquely supporting advanced multimodal and full-duplex behaviors such as dialogue and action turn-taking, defective instruction rejection, speaking-while-acting, context-grounded visual question answering, and action barge-ins. We contend that ELLSA represents a step toward more natural and general interactive intelligence, contributing to the broader pursuit of artificial general intelligence. All data, code and model checkpoints will be released at https://github.com/bytedance/SALMONN/tree/ELLSA.
title End-to-end Listen, Look, Speak and Act
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
Audio and Speech Processing
url https://arxiv.org/abs/2510.16756