IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909898939301888 |
|---|---|
| author | Gao, Yiming Wang, Bin Wei, Chengwei Sun, Shuo Aw, AiTi |
| author_facet | Gao, Yiming Wang, Bin Wei, Chengwei Sun, Shuo Aw, AiTi |
| contents | Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as images or audio. While several recent efforts have investigated instruction-following performance in text and vision-language models, instruction-following in audio-based large language models remains largely unexplored. To bridge this gap, we introduce IFEval-Audio, a novel evaluation dataset designed to assess the ability to follow instructions in an audio LLM. IFEval-Audio contains 280 audio-instruction-answer triples across six diverse dimensions: Content, Capitalization, Symbol, List Structure, Length, and Format. Each example pairs an audio input with a text instruction, requiring the model to generate an output that follows a specified structure. We benchmark state-of-the-art audio LLMs on their ability to follow audio-involved instructions. The dataset is released publicly to support future research in this emerging area. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16774 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models Gao, Yiming Wang, Bin Wei, Chengwei Sun, Shuo Aw, AiTi Computation and Language Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as images or audio. While several recent efforts have investigated instruction-following performance in text and vision-language models, instruction-following in audio-based large language models remains largely unexplored. To bridge this gap, we introduce IFEval-Audio, a novel evaluation dataset designed to assess the ability to follow instructions in an audio LLM. IFEval-Audio contains 280 audio-instruction-answer triples across six diverse dimensions: Content, Capitalization, Symbol, List Structure, Length, and Format. Each example pairs an audio input with a text instruction, requiring the model to generate an output that follows a specified structure. We benchmark state-of-the-art audio LLMs on their ability to follow audio-involved instructions. The dataset is released publicly to support future research in this emerging area. |
| title | IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.16774 |