Evaluating Self-Supervised Speech Models via Text-Based LLMS
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918154777657344 |
|---|---|
| author | Maekaku, Takashi Goto, Keita Tian, Jinchuan Shinohara, Yusuke Watanabe, Shinji |
| author_facet | Maekaku, Takashi Goto, Keita Tian, Jinchuan Shinohara, Yusuke Watanabe, Shinji |
| contents | Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_04463 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evaluating Self-Supervised Speech Models via Text-Based LLMS Maekaku, Takashi Goto, Keita Tian, Jinchuan Shinohara, Yusuke Watanabe, Shinji Sound Audio and Speech Processing Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task. |
| title | Evaluating Self-Supervised Speech Models via Text-Based LLMS |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.04463 |