VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Diatlova, Daria, Balagansky, Nikita, Varlamov, Alexander, Spirin, Egor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911107445161984
author Diatlova, Daria
Balagansky, Nikita
Varlamov, Alexander
Spirin, Egor
author_facet Diatlova, Daria
Balagansky, Nikita
Varlamov, Alexander
Spirin, Egor
contents Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and static feature weighting for all dataset examples. We propose VARAN, a framework that dynamically tailors layer aggregation to individual inputs. By employing layer-specialized probing heads and data-dependent weighting, VARAN adaptively prioritizes layer's features based on input. Evaluations on automatic speech recognition and speech emotion recognition tasks demonstrate VARAN's superior performance, particularly when using the LoRA fine-tuning technique. The framework resolves the trade-off between preserving layer-specific information and enabling flexible feature utilization, advancing efficient adaptation of self-supervised speech representations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
Diatlova, Daria
Balagansky, Nikita
Varlamov, Alexander
Spirin, Egor
Machine Learning
Conventional methods for aggregating layers in fine-tuned self-supervised speech models, such as using the final layer or weighted sum, suffer from information bottlenecks and static feature weighting for all dataset examples. We propose VARAN, a framework that dynamically tailors layer aggregation to individual inputs. By employing layer-specialized probing heads and data-dependent weighting, VARAN adaptively prioritizes layer's features based on input. Evaluations on automatic speech recognition and speech emotion recognition tasks demonstrate VARAN's superior performance, particularly when using the LoRA fine-tuning technique. The framework resolves the trade-off between preserving layer-specific information and enabling flexible feature utilization, advancing efficient adaptation of self-supervised speech representations.
title VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
topic Machine Learning
url https://arxiv.org/abs/2508.12061