Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Poli, Maxime, Chemla, Emmanuel, Dupoux, Emmanuel
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913567710642176
author Poli, Maxime
Chemla, Emmanuel
Dupoux, Emmanuel
author_facet Poli, Maxime
Chemla, Emmanuel
Dupoux, Emmanuel
contents Recent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible. Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. Modeling directly from speech opens up the path to more natural and expressive systems. On the other hand, speech-only systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities. We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
Poli, Maxime
Chemla, Emmanuel
Dupoux, Emmanuel
Computation and Language
Sound
Audio and Speech Processing
Recent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible. Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. Modeling directly from speech opens up the path to more natural and expressive systems. On the other hand, speech-only systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities. We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.
title Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.00025