SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eom, SooHwan, Hasegawa-Johnson, Mark, Yoo, ad Chang D.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916074663968768
author Eom, SooHwan
Hasegawa-Johnson, Mark
Yoo, ad Chang D.
author_facet Eom, SooHwan
Hasegawa-Johnson, Mark
Yoo, ad Chang D.
contents Self-supervised speech representation learning has made significant progress through Siamese networks, which leverage different views of the same input. However, existing methods often require frame-wise alignment between these views, overlooking the broader linguistic context invariance across different speaking styles. We introduce SiamCTC, a framework that integrates Siamese networks with Connectionist Temporal Classification (CTC) to learn speech representations without strict frame-level correspondence. By employing CTC loss to establish flexible, monotonic alignments between differing temporal realizations of the same content, SiamCTC accommodates speed perturbations and other temporal augmentations. This design relaxes frame-wise constraints while preserving temporal coherence and enhancing robustness to speaking-rate variations in downstream tasks. Our experiments demonstrate that SiamCTC leads to more adaptable speech representations, particularly at diverse speaking rates.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
Eom, SooHwan
Hasegawa-Johnson, Mark
Yoo, ad Chang D.
Audio and Speech Processing
Self-supervised speech representation learning has made significant progress through Siamese networks, which leverage different views of the same input. However, existing methods often require frame-wise alignment between these views, overlooking the broader linguistic context invariance across different speaking styles. We introduce SiamCTC, a framework that integrates Siamese networks with Connectionist Temporal Classification (CTC) to learn speech representations without strict frame-level correspondence. By employing CTC loss to establish flexible, monotonic alignments between differing temporal realizations of the same content, SiamCTC accommodates speed perturbations and other temporal augmentations. This design relaxes frame-wise constraints while preserving temporal coherence and enhancing robustness to speaking-rate variations in downstream tasks. Our experiments demonstrate that SiamCTC leads to more adaptable speech representations, particularly at diverse speaking rates.
title SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
topic Audio and Speech Processing
url https://arxiv.org/abs/2606.02220