Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Li-Wei, Higuchi, Takuya, Bai, He, Abdelaziz, Ahmed Hussen, Rudnicky, Alexander, Watanabe, Shinji, Likhomanenko, Tatiana, Theobald, Barry-John, Aldeneh, Zakaria
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913655257300992
author Chen, Li-Wei
Higuchi, Takuya
Bai, He
Abdelaziz, Ahmed Hussen
Rudnicky, Alexander
Watanabe, Shinji
Likhomanenko, Tatiana
Theobald, Barry-John
Aldeneh, Zakaria
author_facet Chen, Li-Wei
Higuchi, Takuya
Bai, He
Abdelaziz, Ahmed Hussen
Rudnicky, Alexander
Watanabe, Shinji
Likhomanenko, Tatiana
Theobald, Barry-John
Aldeneh, Zakaria
contents Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10788
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
Chen, Li-Wei
Higuchi, Takuya
Bai, He
Abdelaziz, Ahmed Hussen
Rudnicky, Alexander
Watanabe, Shinji
Likhomanenko, Tatiana
Theobald, Barry-John
Aldeneh, Zakaria
Audio and Speech Processing
Sound
Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.
title Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.10788