Evaluating Self-Supervised Speech Models via Text-Based LLMS

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maekaku, Takashi, Goto, Keita, Tian, Jinchuan, Shinohara, Yusuke, Watanabe, Shinji
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918154777657344
author Maekaku, Takashi
Goto, Keita
Tian, Jinchuan
Shinohara, Yusuke
Watanabe, Shinji
author_facet Maekaku, Takashi
Goto, Keita
Tian, Jinchuan
Shinohara, Yusuke
Watanabe, Shinji
contents Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Self-Supervised Speech Models via Text-Based LLMS
Maekaku, Takashi
Goto, Keita
Tian, Jinchuan
Shinohara, Yusuke
Watanabe, Shinji
Sound
Audio and Speech Processing
Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task.
title Evaluating Self-Supervised Speech Models via Text-Based LLMS
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.04463