Can Audio Large Language Models Verify Speaker Identity?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ren, Yiming, Xu, Xuenan, Li, Baoxiang, Wang, Shuai, Zhang, Chao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912602442956800
author Ren, Yiming
Xu, Xuenan
Li, Baoxiang
Wang, Shuai
Zhang, Chao
author_facet Ren, Yiming
Xu, Xuenan
Li, Baoxiang
Wang, Shuai
Zhang, Chao
contents This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive zero-shot evaluations on public benchmarks, showing that current ALLMs have limited zero-shot SV capability and often struggle in diverse acoustic conditions. To address this challenge, we perform supervised fine-tuning on speaker verification data. A rule-based hard pair sampling strategy is proposed to construct more challenging training pairs. Lightweight fine-tuning substantially improves the performance, though there is still a gap between ALLMs and conventional models. Then, we extend to text-dependent SV by jointly querying ALLMs to verify speaker identity and spoken content, yielding results competitive with cascaded ASR-SV systems. Our findings demonstrate that with proper adaptation, ALLMs hold substantial potential as a unified model for robust speaker verification systems, while maintaining the general audio understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19755
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Audio Large Language Models Verify Speaker Identity?
Ren, Yiming
Xu, Xuenan
Li, Baoxiang
Wang, Shuai
Zhang, Chao
Sound
Audio and Speech Processing
This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive zero-shot evaluations on public benchmarks, showing that current ALLMs have limited zero-shot SV capability and often struggle in diverse acoustic conditions. To address this challenge, we perform supervised fine-tuning on speaker verification data. A rule-based hard pair sampling strategy is proposed to construct more challenging training pairs. Lightweight fine-tuning substantially improves the performance, though there is still a gap between ALLMs and conventional models. Then, we extend to text-dependent SV by jointly querying ALLMs to verify speaker identity and spoken content, yielding results competitive with cascaded ASR-SV systems. Our findings demonstrate that with proper adaptation, ALLMs hold substantial potential as a unified model for robust speaker verification systems, while maintaining the general audio understanding capabilities.
title Can Audio Large Language Models Verify Speaker Identity?
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.19755