Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Yi, Bai, Jisheng, Xu, Qisheng, Xu, Kele, Dou, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912962502983680
author Su, Yi
Bai, Jisheng
Xu, Qisheng
Xu, Kele
Dou, Yong
author_facet Su, Yi
Bai, Jisheng
Xu, Qisheng
Xu, Kele
Dou, Yong
contents Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage natural language supervision to better handle complex real-world audio scenes with multiple overlapping events. While demonstrating impressive zero-shot and task generalization capabilities, there is still a notable lack of systematic surveys that comprehensively organize and analyze developments. In this paper, we present the first systematic review of ALMs with three main contributions: (1) comprehensive coverage of ALM works across speech, music, and sound from a general audio perspective; (2) a unified taxonomy of ALM foundations, including model architectures and training objectives; (3) establishment of a research landscape capturing mutual promotion and constraints among different research aspects, aiding in summarizing evaluations, limitations, concerns and promising directions. Our review contributes to helping researchers understand the development of existing technologies and future trends, while also providing valuable references for implementation in practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Su, Yi
Bai, Jisheng
Xu, Qisheng
Xu, Kele
Dou, Yong
Sound
Multimedia
Audio and Speech Processing
Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervised approaches that use predefined labels, ALMs leverage natural language supervision to better handle complex real-world audio scenes with multiple overlapping events. While demonstrating impressive zero-shot and task generalization capabilities, there is still a notable lack of systematic surveys that comprehensively organize and analyze developments. In this paper, we present the first systematic review of ALMs with three main contributions: (1) comprehensive coverage of ALM works across speech, music, and sound from a general audio perspective; (2) a unified taxonomy of ALM foundations, including model architectures and training objectives; (3) establishment of a research landscape capturing mutual promotion and constraints among different research aspects, aiding in summarizing evaluations, limitations, concerns and promising directions. Our review contributes to helping researchers understand the development of existing technologies and future trends, while also providing valuable references for implementation in practical applications.
title Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2501.15177