VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908977937252352 |
|---|---|
| author | Lin, Yi-Cheng Hirota, Yusuke Huang, Sung-Feng Lee, Hung-yi |
| author_facet | Lin, Yi-Cheng Hirota, Yusuke Huang, Sung-Feng Lee, Hung-yi |
| contents | Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_17248 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech Lin, Yi-Cheng Hirota, Yusuke Huang, Sung-Feng Lee, Hung-yi Audio and Speech Processing Computation and Language Sound Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes. |
| title | VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2604.17248 |