AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Hoang, Surapaneni, Sidharth, Kalkunte, Akshay, Mehta, Jash, Tiwari, Aman, Bamgbose, Oluwanifemi, Mahajan, Khyati, Shah, Jash, Radhakrishna, Shruthan, Madhusudhan, Sathwik Tejaswi, Yadav, Vikas, Rajeswar, Sai
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918494308663296
author Nguyen, Hoang
Surapaneni, Sidharth
Kalkunte, Akshay
Mehta, Jash
Tiwari, Aman
Bamgbose, Oluwanifemi
Mahajan, Khyati
Shah, Jash
Radhakrishna, Shruthan
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
Rajeswar, Sai
author_facet Nguyen, Hoang
Surapaneni, Sidharth
Kalkunte, Akshay
Mehta, Jash
Tiwari, Aman
Bamgbose, Oluwanifemi
Mahajan, Khyati
Shah, Jash
Radhakrishna, Shruthan
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
Rajeswar, Sai
contents Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized toolkits that limit fair comparison and systematic assessment. Existing evaluation frameworks exhibit three critical limitations: (1) slow and inefficient processing pipeline that bottlenecks large-scale studies, (2) inadequate multi-turn dialogue support, leaving fundamental questions about cross-turn context integration and performance dynamics over extended conversations in LALMs unanswered; and (3) the absence of unified and scalable evaluation framework capable of keeping pace with the rapid growth of both LALMs and audio benchmarks. To address these issues, we introduce AU-Harness, an efficient and comprehensive evaluation framework for LALMs. Our system achieves a speedup of up to 151% over existing evaluation toolkits through optimized batch processing and parallel execution, enabling large-scale evaluations previously considered impractical. We provide standardized prompting protocols and flexible configurations for fair model comparison across diverse scenarios. AU-Harness unlocks a range of in-depth analyses difficult to conduct without a unified foundation, including multi-turn dialogue dynamics, enabling the study of true audio reasoning capabilities in existing LALMs. AU-Harness provides both practical evaluation tools and insights into model limitations, advancing systematic LALM development.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
Nguyen, Hoang
Surapaneni, Sidharth
Kalkunte, Akshay
Mehta, Jash
Tiwari, Aman
Bamgbose, Oluwanifemi
Mahajan, Khyati
Shah, Jash
Radhakrishna, Shruthan
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
Rajeswar, Sai
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized toolkits that limit fair comparison and systematic assessment. Existing evaluation frameworks exhibit three critical limitations: (1) slow and inefficient processing pipeline that bottlenecks large-scale studies, (2) inadequate multi-turn dialogue support, leaving fundamental questions about cross-turn context integration and performance dynamics over extended conversations in LALMs unanswered; and (3) the absence of unified and scalable evaluation framework capable of keeping pace with the rapid growth of both LALMs and audio benchmarks. To address these issues, we introduce AU-Harness, an efficient and comprehensive evaluation framework for LALMs. Our system achieves a speedup of up to 151% over existing evaluation toolkits through optimized batch processing and parallel execution, enabling large-scale evaluations previously considered impractical. We provide standardized prompting protocols and flexible configurations for fair model comparison across diverse scenarios. AU-Harness unlocks a range of in-depth analyses difficult to conduct without a unified foundation, including multi-turn dialogue dynamics, enabling the study of true audio reasoning capabilities in existing LALMs. AU-Harness provides both practical evaluation tools and insights into model limitations, advancing systematic LALM development.
title AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.08031