Audio-Visual Intelligence in Large Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qin, You, Liu, Kai, Wu, Shengqiong, Wang, Kai, Deng, Shijian, Tian, Yapeng, Xiao, Junbin, Xing, Yazhou, Ma, Yinghao, Li, Bobo, Zimmermann, Roger, Cui, Lei, Wei, Furu, Luo, Jiebo, Fei, Hao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910192501784576
author Qin, You
Liu, Kai
Wu, Shengqiong
Wang, Kai
Deng, Shijian
Tian, Yapeng
Xiao, Junbin
Xing, Yazhou
Ma, Yinghao
Li, Bobo
Zimmermann, Roger
Cui, Lei
Wei, Furu
Luo, Jiebo
Fei, Hao
author_facet Qin, You
Liu, Kai
Wu, Shengqiong
Wang, Kai
Deng, Shijian
Tian, Yapeng
Xiao, Junbin
Xing, Yazhou
Ma, Yinghao
Li, Bobo
Zimmermann, Roger
Cui, Lei
Wei, Furu
Luo, Jiebo
Fei, Hao
contents Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo-3, highlight the growing industrial and academic focus on unified audio-vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio-driven video synthesis, video-to-audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large-scale AVI.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04045
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Audio-Visual Intelligence in Large Foundation Models
Qin, You
Liu, Kai
Wu, Shengqiong
Wang, Kai
Deng, Shijian
Tian, Yapeng
Xiao, Junbin
Xing, Yazhou
Ma, Yinghao
Li, Bobo
Zimmermann, Roger
Cui, Lei
Wei, Furu
Luo, Jiebo
Fei, Hao
Computer Vision and Pattern Recognition
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo-3, highlight the growing industrial and academic focus on unified audio-vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio-driven video synthesis, video-to-audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large-scale AVI.
title Audio-Visual Intelligence in Large Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.04045