Saved in:
| Main Authors: | Mayer, Paul, Lux, Florian, Pérez-González-de-Martos, Alejandro, Elizarova, Angelina, Vanderlyn, Lindsey, Väth, Dirk, Vu, Ngoc Thang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.00227 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Controlling Emotion in Text-to-Speech with Natural Language Prompts
by: Bott, Thomas, et al.
Published: (2024)
by: Bott, Thomas, et al.
Published: (2024)
Probing the Feasibility of Multilingual Speaker Anonymization
by: Meyer, Sarina, et al.
Published: (2024)
by: Meyer, Sarina, et al.
Published: (2024)
Investigating the effect of Mental Models in User Interaction with an Adaptive Dialog Agent
by: Vanderlyn, Lindsey, et al.
Published: (2024)
by: Vanderlyn, Lindsey, et al.
Published: (2024)
High-Resolution Speech Restoration with Latent Diffusion Model
by: Dhyani, Tushar, et al.
Published: (2024)
by: Dhyani, Tushar, et al.
Published: (2024)
First Steps Towards Voice Anonymization for Code-Switching Speech
by: Meyer, Sarina, et al.
Published: (2025)
by: Meyer, Sarina, et al.
Published: (2025)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
by: Lux, Florian, et al.
Published: (2024)
by: Lux, Florian, et al.
Published: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
by: Denisov, Pavel, et al.
Published: (2024)
by: Denisov, Pavel, et al.
Published: (2024)
Towards a Zero-Data, Controllable, Adaptive Dialog System
by: Väth, Dirk, et al.
Published: (2024)
by: Väth, Dirk, et al.
Published: (2024)
Conversational Tree Search: A New Hybrid Dialog Task
by: Väth, Dirk, et al.
Published: (2023)
by: Väth, Dirk, et al.
Published: (2023)
Use Cases for Voice Anonymization
by: Meyer, Sarina, et al.
Published: (2025)
by: Meyer, Sarina, et al.
Published: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
by: Oh, Hyung-Seok, et al.
Published: (2023)
by: Oh, Hyung-Seok, et al.
Published: (2023)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
by: Li, Chia-Yu, et al.
Published: (2024)
by: Li, Chia-Yu, et al.
Published: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
by: Jiang, Yuepeng, et al.
Published: (2024)
by: Jiang, Yuepeng, et al.
Published: (2024)
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
by: Qian, Kaizhi, et al.
Published: (2025)
by: Qian, Kaizhi, et al.
Published: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
by: Du, Chenpeng, et al.
Published: (2021)
by: Du, Chenpeng, et al.
Published: (2021)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
by: Koriyama, Tomoki
Published: (2025)
by: Koriyama, Tomoki
Published: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
by: He, Xiangheng, et al.
Published: (2024)
by: He, Xiangheng, et al.
Published: (2024)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
by: Girish, et al.
Published: (2026)
by: Girish, et al.
Published: (2026)
The Risks and Detection of Overestimated Privacy Protection in Voice Anonymisation
by: Panariello, Michele, et al.
Published: (2025)
by: Panariello, Michele, et al.
Published: (2025)
Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated Approach
by: Chan, Cedric, et al.
Published: (2025)
by: Chan, Cedric, et al.
Published: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
by: Onda, Kentaro, et al.
Published: (2025)
by: Onda, Kentaro, et al.
Published: (2025)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
by: Karapiperis, Sotirios, et al.
Published: (2024)
by: Karapiperis, Sotirios, et al.
Published: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
by: Han, Wooseok, et al.
Published: (2024)
by: Han, Wooseok, et al.
Published: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
Privacy-preserving Prosody Representation Learning
by: Everson, Kevin, et al.
Published: (2026)
by: Everson, Kevin, et al.
Published: (2026)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
by: Zhao, Junchuan, et al.
Published: (2026)
by: Zhao, Junchuan, et al.
Published: (2026)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
by: Borodin, Kirill, et al.
Published: (2025)
by: Borodin, Kirill, et al.
Published: (2025)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
by: Al-Radhi, Mohammed Salah, et al.
Published: (2025)
by: Al-Radhi, Mohammed Salah, et al.
Published: (2025)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
by: Mondal, Anindita, et al.
Published: (2024)
by: Mondal, Anindita, et al.
Published: (2024)
Prosody Analysis of Audiobooks
by: Pethe, Charuta, et al.
Published: (2023)
by: Pethe, Charuta, et al.
Published: (2023)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
by: Kheir, Yassine El, et al.
Published: (2026)
by: Kheir, Yassine El, et al.
Published: (2026)
AS-Speech: Adaptive Style For Speech Synthesis
by: Li, Zhipeng, et al.
Published: (2024)
by: Li, Zhipeng, et al.
Published: (2024)
An Attribute Interpolation Method in Speech Synthesis by Model Merging
by: Murata, Masato, et al.
Published: (2024)
by: Murata, Masato, et al.
Published: (2024)
PRESENT: Zero-Shot Text-to-Prosody Control
by: Lam, Perry, et al.
Published: (2024)
by: Lam, Perry, et al.
Published: (2024)
Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
by: Mahapatra, Aurosweta, et al.
Published: (2025)
by: Mahapatra, Aurosweta, et al.
Published: (2025)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation
by: Koch, Timo K., et al.
Published: (2026)
by: Koch, Timo K., et al.
Published: (2026)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
by: Wu, Ya-Tse, et al.
Published: (2026)
by: Wu, Ya-Tse, et al.
Published: (2026)
Similar Items
-
Controlling Emotion in Text-to-Speech with Natural Language Prompts
by: Bott, Thomas, et al.
Published: (2024) -
Probing the Feasibility of Multilingual Speaker Anonymization
by: Meyer, Sarina, et al.
Published: (2024) -
Investigating the effect of Mental Models in User Interaction with an Adaptive Dialog Agent
by: Vanderlyn, Lindsey, et al.
Published: (2024) -
High-Resolution Speech Restoration with Latent Diffusion Model
by: Dhyani, Tushar, et al.
Published: (2024) -
First Steps Towards Voice Anonymization for Code-Switching Speech
by: Meyer, Sarina, et al.
Published: (2025)