Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Thebaud, Thomas, Wang, Yuzhe, Moro-Velazquez, Laureano, Villalba-Lopez, Jesus, Dehak, Najim
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914386384257024
author Thebaud, Thomas
Wang, Yuzhe
Moro-Velazquez, Laureano
Villalba-Lopez, Jesus
Dehak, Najim
author_facet Thebaud, Thomas
Wang, Yuzhe
Moro-Velazquez, Laureano
Villalba-Lopez, Jesus
Dehak, Najim
contents Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity. First, we propose a model-agnostic scoring protocol that produces continuous verification scores for both API-only and open-weight models, using confidence scores or log-likelihood ratios from the Yes/No token probabilities. Using this protocol, we benchmark recent speech-aware LLMs and observe weak speaker discrimination (EERs above 20% on VoxCeleb1). Second, we introduce a lightweight augmentation that equips an LLM with ASV capability by injecting frozen ECAPA-TDNN speaker embeddings through a learned projection and training only LoRA adapters. On TinyLLaMA-1.1B, the resulting ECAPA-LLM achieves 1.03% EER on VoxCeleb1-E, approaching a dedicated speaker verification system while preserving a natural-language interface.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10827
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
Thebaud, Thomas
Wang, Yuzhe
Moro-Velazquez, Laureano
Villalba-Lopez, Jesus
Dehak, Najim
Sound
Artificial Intelligence
Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity. First, we propose a model-agnostic scoring protocol that produces continuous verification scores for both API-only and open-weight models, using confidence scores or log-likelihood ratios from the Yes/No token probabilities. Using this protocol, we benchmark recent speech-aware LLMs and observe weak speaker discrimination (EERs above 20% on VoxCeleb1). Second, we introduce a lightweight augmentation that equips an LLM with ASV capability by injecting frozen ECAPA-TDNN speaker embeddings through a learned projection and training only LoRA adapters. On TinyLLaMA-1.1B, the resulting ECAPA-LLM achieves 1.03% EER on VoxCeleb1-E, approaching a dedicated speaker verification system while preserving a natural-language interface.
title Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2603.10827