Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zaiem, Salah, Kemiche, Youcef, Parcollet, Titouan, Essid, Slim, Ravanelli, Mirco
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917594686029824
author Zaiem, Salah
Kemiche, Youcef
Parcollet, Titouan
Essid, Slim
Ravanelli, Mirco
author_facet Zaiem, Salah
Kemiche, Youcef
Parcollet, Titouan
Essid, Slim
Ravanelli, Mirco
contents Self-supervised learning (SSL) leverages large datasets of unlabeled speech to reach impressive performance with reduced amounts of annotated data. The high number of proposed approaches fostered the emergence of comprehensive benchmarks that evaluate their performance on a set of downstream tasks exploring various aspects of the speech signal. However, while the number of considered tasks has been growing, most proposals rely upon a single downstream architecture that maps the frozen SSL representations to the task labels. This study examines how benchmarking results are affected by changes in the probing head architecture. Interestingly, we found that altering the downstream architecture structure leads to significant fluctuations in the performance ranking of the evaluated models. Against common practices in speech SSL benchmarking, we evaluate larger-capacity probing heads, showing their impact on performance, inference costs, generalization and multi-level feature exploitation.
format Preprint
id arxiv_https___arxiv_org_abs_2308_14456
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
Zaiem, Salah
Kemiche, Youcef
Parcollet, Titouan
Essid, Slim
Ravanelli, Mirco
Audio and Speech Processing
Machine Learning
Sound
Signal Processing
Self-supervised learning (SSL) leverages large datasets of unlabeled speech to reach impressive performance with reduced amounts of annotated data. The high number of proposed approaches fostered the emergence of comprehensive benchmarks that evaluate their performance on a set of downstream tasks exploring various aspects of the speech signal. However, while the number of considered tasks has been growing, most proposals rely upon a single downstream architecture that maps the frozen SSL representations to the task labels. This study examines how benchmarking results are affected by changes in the probing head architecture. Interestingly, we found that altering the downstream architecture structure leads to significant fluctuations in the performance ranking of the evaluated models. Against common practices in speech SSL benchmarking, we evaluate larger-capacity probing heads, showing their impact on performance, inference costs, generalization and multi-level feature exploitation.
title Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
topic Audio and Speech Processing
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2308.14456