Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chorowski, Jan, Ciesielski, Grzegorz, Dzikowski, Jarosław, Łańcucki, Adrian, Marxer, Ricard, Opala, Mateusz, Pusz, Piotr, Rychlikowski, Paweł, Stypułkowski, Michał
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910600082227200
author Chorowski, Jan
Ciesielski, Grzegorz
Dzikowski, Jarosław
Łańcucki, Adrian
Marxer, Ricard
Opala, Mateusz
Pusz, Piotr
Rychlikowski, Paweł
Stypułkowski, Michał
author_facet Chorowski, Jan
Ciesielski, Grzegorz
Dzikowski, Jarosław
Łańcucki, Adrian
Marxer, Ricard
Opala, Mateusz
Pusz, Piotr
Rychlikowski, Paweł
Stypułkowski, Michał
contents We present a number of low-resource approaches to the tasks of the Zero Resource Speech Challenge 2021. We build on the unsupervised representations of speech proposed by the organizers as a baseline, derived from CPC and clustered with the k-means algorithm. We demonstrate that simple methods of refining those representations can narrow the gap, or even improve upon the solutions which use a high computational budget. The results lead to the conclusion that the CPC-derived representations are still too noisy for training language models, but stable enough for simpler forms of pattern matching and retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2106_11603
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
Chorowski, Jan
Ciesielski, Grzegorz
Dzikowski, Jarosław
Łańcucki, Adrian
Marxer, Ricard
Opala, Mateusz
Pusz, Piotr
Rychlikowski, Paweł
Stypułkowski, Michał
Machine Learning
Sound
Audio and Speech Processing
We present a number of low-resource approaches to the tasks of the Zero Resource Speech Challenge 2021. We build on the unsupervised representations of speech proposed by the organizers as a baseline, derived from CPC and clustered with the k-means algorithm. We demonstrate that simple methods of refining those representations can narrow the gap, or even improve upon the solutions which use a high computational budget. The results lead to the conclusion that the CPC-derived representations are still too noisy for training language models, but stable enough for simpler forms of pattern matching and retrieval.
title Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2106.11603