CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Deng, Ruifan, Gong, Yitian, Gao, Qinghui, Jin, Luozhijie, Cheng, Qinyuan, Fei, Zhaoye, Li, Shimin, Qiu, Xipeng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915468087918592
author Deng, Ruifan
Gong, Yitian
Gao, Qinghui
Jin, Luozhijie
Cheng, Qinyuan
Fei, Zhaoye
Li, Shimin
Qiu, Xipeng
author_facet Deng, Ruifan
Gong, Yitian
Gao, Qinghui
Jin, Luozhijie
Cheng, Qinyuan
Fei, Zhaoye
Li, Shimin
Qiu, Xipeng
contents With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into text-based LLMs. Current audio codec captures two types of information: acoustic and semantic. As audio codec is applied to diverse scenarios in speech language model , it needs to model increasingly complex information and adapt to varied contexts, such as scenarios with multiple speakers, background noise, or richer paralinguistic information. However, existing codec's own evaluation has been limited by simplistic metrics and scenarios, and existing benchmarks for audio codec are not designed for complex application scenarios, which limits the assessment performance on complex datasets for acoustic and semantic capabilities. We introduce CodecBench, a comprehensive evaluation dataset to assess audio codec performance from both acoustic and semantic perspectives across four data domains. Through this benchmark, we aim to identify current limitations, highlight future research directions, and foster advances in the development of audio codec. The codes are available at https://github.com/RayYuki/CodecBench.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20660
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
Deng, Ruifan
Gong, Yitian
Gao, Qinghui
Jin, Luozhijie
Cheng, Qinyuan
Fei, Zhaoye
Li, Shimin
Qiu, Xipeng
Audio and Speech Processing
Sound
With the rise of multimodal large language models (LLMs), audio codec plays an increasingly vital role in encoding audio into discrete tokens, enabling integration of audio into text-based LLMs. Current audio codec captures two types of information: acoustic and semantic. As audio codec is applied to diverse scenarios in speech language model , it needs to model increasingly complex information and adapt to varied contexts, such as scenarios with multiple speakers, background noise, or richer paralinguistic information. However, existing codec's own evaluation has been limited by simplistic metrics and scenarios, and existing benchmarks for audio codec are not designed for complex application scenarios, which limits the assessment performance on complex datasets for acoustic and semantic capabilities. We introduce CodecBench, a comprehensive evaluation dataset to assess audio codec performance from both acoustic and semantic perspectives across four data domains. Through this benchmark, we aim to identify current limitations, highlight future research directions, and foster advances in the development of audio codec. The codes are available at https://github.com/RayYuki/CodecBench.
title CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2508.20660