ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Eren, Eray, Liu, Qingju, Kim, Hyeongwoo, Garrido, Pablo, Alwan, Abeer
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912535817486336
author Eren, Eray
Liu, Qingju
Kim, Hyeongwoo
Garrido, Pablo
Alwan, Abeer
author_facet Eren, Eray
Liu, Qingju
Kim, Hyeongwoo
Garrido, Pablo
Alwan, Abeer
contents Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks such as TTS. The ProMode encoder takes as input acoustic features and time-aligned textual content, both are partially masked, and obtains a fixed-length latent prosodic embedding. The decoder predicts acoustics in the masked region using both the encoded prosody input and unmasked textual content. Trained on the GigaSpeech dataset, we compare our method with state-of-the-art style encoders. For F0 and energy predictions, we show consistent improvements for our model at different levels of granularity. We also integrate these predicted prosodic features into a TTS system and conduct perceptual tests, which show higher prosody preference compared to the baselines, demonstrating the model's potential in tasks where prosody modeling is important.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
Eren, Eray
Liu, Qingju
Kim, Hyeongwoo
Garrido, Pablo
Alwan, Abeer
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks such as TTS. The ProMode encoder takes as input acoustic features and time-aligned textual content, both are partially masked, and obtains a fixed-length latent prosodic embedding. The decoder predicts acoustics in the masked region using both the encoded prosody input and unmasked textual content. Trained on the GigaSpeech dataset, we compare our method with state-of-the-art style encoders. For F0 and energy predictions, we show consistent improvements for our model at different levels of granularity. We also integrate these predicted prosodic features into a TTS system and conduct perceptual tests, which show higher prosody preference compared to the baselines, demonstrating the model's potential in tasks where prosody modeling is important.
title ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2508.09389