MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tam, Zhi Rui, Chen, Yun-Nung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908640151076864
author Tam, Zhi Rui
Chen, Yun-Nung
author_facet Tam, Zhi Rui
Chen, Yun-Nung
contents As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in audio. We evaluated these models on 170 clinical cases, each synthesized into speech from 36 distinct voice profiles spanning variations in age, gender, and emotion. Our findings reveal a severe modality bias: surgical recommendations for audio inputs varied by as much as 35% compared to identical text-based inputs, with one model providing 80% fewer recommendations. Further analysis uncovered age disparities of up to 12% between young and elderly voices, which persisted in most models despite chain-of-thought prompting. While explicit reasoning successfully eliminated gender bias, the impact of emotion was not detected due to poor recognition performance. These results demonstrate that audio LLMs are susceptible to making clinical decisions based on a patient's voice characteristics rather than medical evidence, a flaw that risks perpetuating healthcare disparities. We conclude that bias-aware architectures are essential and urgently needed before the clinical deployment of these models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
Tam, Zhi Rui
Chen, Yun-Nung
Computation and Language
Audio and Speech Processing
As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in audio. We evaluated these models on 170 clinical cases, each synthesized into speech from 36 distinct voice profiles spanning variations in age, gender, and emotion. Our findings reveal a severe modality bias: surgical recommendations for audio inputs varied by as much as 35% compared to identical text-based inputs, with one model providing 80% fewer recommendations. Further analysis uncovered age disparities of up to 12% between young and elderly voices, which persisted in most models despite chain-of-thought prompting. While explicit reasoning successfully eliminated gender bias, the impact of emotion was not detected due to poor recognition performance. These results demonstrate that audio LLMs are susceptible to making clinical decisions based on a patient's voice characteristics rather than medical evidence, a flaw that risks perpetuating healthcare disparities. We conclude that bias-aware architectures are essential and urgently needed before the clinical deployment of these models.
title MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2511.06592