VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Yi-Cheng, Hirota, Yusuke, Huang, Sung-Feng, Lee, Hung-yi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908977937252352
author Lin, Yi-Cheng
Hirota, Yusuke
Huang, Sung-Feng
Lee, Hung-yi
author_facet Lin, Yi-Cheng
Hirota, Yusuke
Huang, Sung-Feng
Lee, Hung-yi
contents Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
Lin, Yi-Cheng
Hirota, Yusuke
Huang, Sung-Feng
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Sound
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using real-world human recordings. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 11 state-of-the-art LALMs reveals systematic biases in realistic scenarios. We find that gender cues often trigger larger distributional shifts than accent cues, indicating that current LALMs reproduce social stereotypes.
title VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2604.17248