The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Luxi, Qi, Xiangyu, Liao, Michel, Cheong, Inyoung, Mittal, Prateek, Chen, Danqi, Henderson, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915484943777792
author He, Luxi
Qi, Xiangyu
Liao, Michel
Cheong, Inyoung
Mittal, Prateek
Chen, Danqi
Henderson, Peter
author_facet He, Luxi
Qi, Xiangyu
Liao, Michel
Cheong, Inyoung
Mittal, Prateek
Chen, Danqi
Henderson, Peter
contents The latest Audio Language Models (Audio LMs) process speech directly instead of relying on a separate transcription step. This shift preserves detailed information, such as intonation or the presence of multiple speakers, that would otherwise be lost in transcription. However, it also introduces new safety risks, including the potential misuse of speaker identity cues and other sensitive vocal attributes, which could have legal implications. In this paper, we urge a closer examination of how these models are built and deployed. Our experiments show that end-to-end modeling, compared with cascaded pipelines, creates socio-technical safety risks such as identity inference, biased decision-making, and emotion detection. This raises concerns about whether Audio LMs store voiceprints and function in ways that create uncertainty under existing legal regimes. We then argue that the Principle of Least Privilege should be considered to guide the development and deployment of these models. Specifically, evaluations should assess (1) the privacy and safety risks associated with end-to-end modeling; and (2) the appropriate scope of information access. Finally, we highlight related gaps in current audio LM benchmarks and identify key open research questions, both technical and policy-related, that must be addressed to enable the responsible deployment of end-to-end Audio LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16833
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
He, Luxi
Qi, Xiangyu
Liao, Michel
Cheong, Inyoung
Mittal, Prateek
Chen, Danqi
Henderson, Peter
Sound
Artificial Intelligence
Computation and Language
Computers and Society
Audio and Speech Processing
The latest Audio Language Models (Audio LMs) process speech directly instead of relying on a separate transcription step. This shift preserves detailed information, such as intonation or the presence of multiple speakers, that would otherwise be lost in transcription. However, it also introduces new safety risks, including the potential misuse of speaker identity cues and other sensitive vocal attributes, which could have legal implications. In this paper, we urge a closer examination of how these models are built and deployed. Our experiments show that end-to-end modeling, compared with cascaded pipelines, creates socio-technical safety risks such as identity inference, biased decision-making, and emotion detection. This raises concerns about whether Audio LMs store voiceprints and function in ways that create uncertainty under existing legal regimes. We then argue that the Principle of Least Privilege should be considered to guide the development and deployment of these models. Specifically, evaluations should assess (1) the privacy and safety risks associated with end-to-end modeling; and (2) the appropriate scope of information access. Finally, we highlight related gaps in current audio LM benchmarks and identify key open research questions, both technical and policy-related, that must be addressed to enable the responsible deployment of end-to-end Audio LMs.
title The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
topic Sound
Artificial Intelligence
Computation and Language
Computers and Society
Audio and Speech Processing
url https://arxiv.org/abs/2503.16833