Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cuervo, Santiago, Grabias, Maciej, Chorowski, Jan, Ciesielski, Grzegorz, Łańcucki, Adrian, Rychlikowski, Paweł, Marxer, Ricard
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929497259900928
author Cuervo, Santiago
Grabias, Maciej
Chorowski, Jan
Ciesielski, Grzegorz
Łańcucki, Adrian
Rychlikowski, Paweł
Marxer, Ricard
author_facet Cuervo, Santiago
Grabias, Maciej
Chorowski, Jan
Ciesielski, Grzegorz
Łańcucki, Adrian
Rychlikowski, Paweł
Marxer, Ricard
contents We investigate the performance on phoneme categorization and phoneme and word segmentation of several self-supervised learning (SSL) methods based on Contrastive Predictive Coding (CPC). Our experiments show that with the existing algorithms there is a trade off between categorization and segmentation performance. We investigate the source of this conflict and conclude that the use of context building networks, albeit necessary for superior performance on categorization tasks, harms segmentation performance by causing a temporal shift on the learned representations. Aiming to bridge this gap, we take inspiration from the leading approach on segmentation, which simultaneously models the speech signal at the frame and phoneme level, and incorporate multi-level modelling into Aligned CPC (ACPC), a variation of CPC which exhibits the best performance on categorization tasks. Our multi-level ACPC (mACPC) improves in all categorization metrics and achieves state-of-the-art performance in word segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2110_15909
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
Cuervo, Santiago
Grabias, Maciej
Chorowski, Jan
Ciesielski, Grzegorz
Łańcucki, Adrian
Rychlikowski, Paweł
Marxer, Ricard
Machine Learning
Sound
Audio and Speech Processing
We investigate the performance on phoneme categorization and phoneme and word segmentation of several self-supervised learning (SSL) methods based on Contrastive Predictive Coding (CPC). Our experiments show that with the existing algorithms there is a trade off between categorization and segmentation performance. We investigate the source of this conflict and conclude that the use of context building networks, albeit necessary for superior performance on categorization tasks, harms segmentation performance by causing a temporal shift on the learned representations. Aiming to bridge this gap, we take inspiration from the leading approach on segmentation, which simultaneously models the speech signal at the frame and phoneme level, and incorporate multi-level modelling into Aligned CPC (ACPC), a variation of CPC which exhibits the best performance on categorization tasks. Our multi-level ACPC (mACPC) improves in all categorization metrics and achieves state-of-the-art performance in word segmentation.
title Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2110.15909