Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914386384257024 |
|---|---|
| author | Thebaud, Thomas Wang, Yuzhe Moro-Velazquez, Laureano Villalba-Lopez, Jesus Dehak, Najim |
| author_facet | Thebaud, Thomas Wang, Yuzhe Moro-Velazquez, Laureano Villalba-Lopez, Jesus Dehak, Najim |
| contents | Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity. First, we propose a model-agnostic scoring protocol that produces continuous verification scores for both API-only and open-weight models, using confidence scores or log-likelihood ratios from the Yes/No token probabilities. Using this protocol, we benchmark recent speech-aware LLMs and observe weak speaker discrimination (EERs above 20% on VoxCeleb1). Second, we introduce a lightweight augmentation that equips an LLM with ASV capability by injecting frozen ECAPA-TDNN speaker embeddings through a learned projection and training only LoRA adapters. On TinyLLaMA-1.1B, the resulting ECAPA-LLM achieves 1.03% EER on VoxCeleb1-E, approaching a dedicated speaker verification system while preserving a natural-language interface. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_10827 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation Thebaud, Thomas Wang, Yuzhe Moro-Velazquez, Laureano Villalba-Lopez, Jesus Dehak, Najim Sound Artificial Intelligence Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the speaker's gender, leaving it unclear whether they encode speaker identity. First, we propose a model-agnostic scoring protocol that produces continuous verification scores for both API-only and open-weight models, using confidence scores or log-likelihood ratios from the Yes/No token probabilities. Using this protocol, we benchmark recent speech-aware LLMs and observe weak speaker discrimination (EERs above 20% on VoxCeleb1). Second, we introduce a lightweight augmentation that equips an LLM with ASV capability by injecting frozen ECAPA-TDNN speaker embeddings through a learned projection and training only LoRA adapters. On TinyLLaMA-1.1B, the resulting ECAPA-LLM achieves 1.03% EER on VoxCeleb1-E, approaching a dedicated speaker verification system while preserving a natural-language interface. |
| title | Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2603.10827 |