Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sanders, Nicholas, Li, Yuanchao, Richmond, Korin, King, Simon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912385963393024
author Sanders, Nicholas
Li, Yuanchao
Richmond, Korin
King, Simon
author_facet Sanders, Nicholas
Li, Yuanchao
Richmond, Korin
King, Simon
contents Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence). While increasing codebook size mitigates some loss, it inefficiently raises bitrates. We propose Segmentation-Variant Codebooks (SVCs), which quantize speech at distinct linguistic units (frame, phone, word, utterance), factorizing it into multiple streams of segment-specific discrete features. Our results show that SVCs are significantly more effective at preserving prosodic and paralinguistic information across probing tasks. Additionally, we find that pooling before rather than after discretization better retains segment-level information. Resynthesis experiments further confirm improved style realization and slightly improved quality while preserving intelligibility.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Sanders, Nicholas
Li, Yuanchao
Richmond, Korin
King, Simon
Audio and Speech Processing
Computation and Language
Sound
Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence). While increasing codebook size mitigates some loss, it inefficiently raises bitrates. We propose Segmentation-Variant Codebooks (SVCs), which quantize speech at distinct linguistic units (frame, phone, word, utterance), factorizing it into multiple streams of segment-specific discrete features. Our results show that SVCs are significantly more effective at preserving prosodic and paralinguistic information across probing tasks. Additionally, we find that pooling before rather than after discretization better retains segment-level information. Resynthesis experiments further confirm improved style realization and slightly improved quality while preserving intelligibility.
title Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2505.15667