UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Qundong, Zhou, Jie, Lin, Biyuan, Cui, Junbo, Zeng, Guoyang, Zhou, Yixuan, Wang, Ziyang, Liu, Xin, Luo, Zhen, Wang, Yudong, Liu, Zhiyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917183143018496
author Shi, Qundong
Zhou, Jie
Lin, Biyuan
Cui, Junbo
Zeng, Guoyang
Zhou, Yixuan
Wang, Ziyang
Liu, Xin
Luo, Zhen
Wang, Yudong
Liu, Zhiyuan
author_facet Shi, Qundong
Zhou, Jie
Lin, Biyuan
Cui, Junbo
Zeng, Guoyang
Zhou, Yixuan
Wang, Ziyang
Liu, Xin
Luo, Zhen
Wang, Yudong
Liu, Zhiyuan
contents The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio generation. Current audio evaluation faces three major challenges: (1) audio evaluation lacks a unified framework, with datasets and code scattered across various sources, hindering fair and efficient cross-model comparison;(2) audio codecs, as a key component of audio foundation models, lack a widely accepted and holistic evaluation methodology; (3) existing speech benchmarks are heavily reliant on English, making it challenging to objectively assess models' performance on Chinese. To address the first issue, we introduce UltraEval-Audio, a unified evaluation framework for audio foundation models, specifically designed for both audio understanding and generation tasks. UltraEval-Audio features a modular architecture, supporting 10 languages and 14 core task categories, while seamlessly integrating 24 mainstream models and 36 authoritative benchmarks. To enhance research efficiency, the framework provides a one-command evaluation feature, accompanied by real-time public leaderboards. For the second challenge, UltraEval-Audio adopts a novel comprehensive evaluation scheme for audio codecs, evaluating performance across three key dimensions: semantic accuracy, timbre fidelity, and acoustic quality. To address the third issue, we propose two new Chinese benchmarks, SpeechCMMLU and SpeechHSK, designed to assess Chinese knowledge proficiency and language fluency. We wish that UltraEval-Audio will provide both academia and industry with a transparent, efficient, and fair platform for comparison of audio models. Our code, benchmarks, and leaderboards are available at https://github.com/OpenBMB/UltraEval-Audio.
format Preprint
id arxiv_https___arxiv_org_abs_2601_01373
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
Shi, Qundong
Zhou, Jie
Lin, Biyuan
Cui, Junbo
Zeng, Guoyang
Zhou, Yixuan
Wang, Ziyang
Liu, Xin
Luo, Zhen
Wang, Yudong
Liu, Zhiyuan
Sound
Artificial Intelligence
Audio and Speech Processing
The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particularly in audio generation. Current audio evaluation faces three major challenges: (1) audio evaluation lacks a unified framework, with datasets and code scattered across various sources, hindering fair and efficient cross-model comparison;(2) audio codecs, as a key component of audio foundation models, lack a widely accepted and holistic evaluation methodology; (3) existing speech benchmarks are heavily reliant on English, making it challenging to objectively assess models' performance on Chinese. To address the first issue, we introduce UltraEval-Audio, a unified evaluation framework for audio foundation models, specifically designed for both audio understanding and generation tasks. UltraEval-Audio features a modular architecture, supporting 10 languages and 14 core task categories, while seamlessly integrating 24 mainstream models and 36 authoritative benchmarks. To enhance research efficiency, the framework provides a one-command evaluation feature, accompanied by real-time public leaderboards. For the second challenge, UltraEval-Audio adopts a novel comprehensive evaluation scheme for audio codecs, evaluating performance across three key dimensions: semantic accuracy, timbre fidelity, and acoustic quality. To address the third issue, we propose two new Chinese benchmarks, SpeechCMMLU and SpeechHSK, designed to assess Chinese knowledge proficiency and language fluency. We wish that UltraEval-Audio will provide both academia and industry with a transparent, efficient, and fair platform for comparison of audio models. Our code, benchmarks, and leaderboards are available at https://github.com/OpenBMB/UltraEval-Audio.
title UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2601.01373