Speech Synthesis along Perceptual Voice Quality Dimensions
Fuente:
arXiv
Saved in:
| Main Authors: | Rautenberg, Frederik, Kuhlmann, Michael, Seebauer, Fritz, Wiechmann, Jana, Wagner, Petra, Haeb-Umbach, Reinhold |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
by: Rautenberg, Frederik, et al.
Published: (2026)
by: Rautenberg, Frederik, et al.
Published: (2026)
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
by: Rautenberg, Frederik, et al.
Published: (2025)
by: Rautenberg, Frederik, et al.
Published: (2025)
Towards Frame-level Quality Predictions of Synthetic Speech
by: Kuhlmann, Michael, et al.
Published: (2025)
by: Kuhlmann, Michael, et al.
Published: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
by: Xie, Yuying, et al.
Published: (2024)
by: Xie, Yuying, et al.
Published: (2024)
Diminishing Domain Mismatch for DNN-Based Acoustic Distance Estimation via Stochastic Room Reverberation Models
by: Gburrek, Tobias, et al.
Published: (2024)
by: Gburrek, Tobias, et al.
Published: (2024)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
by: von Neumann, Thilo, et al.
Published: (2023)
by: von Neumann, Thilo, et al.
Published: (2023)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
by: Boeddeker, Christoph, et al.
Published: (2023)
by: Boeddeker, Christoph, et al.
Published: (2023)
30+ Years of Source Separation Research: Achievements and Future Challenges
by: Araki, Shoko, et al.
Published: (2025)
by: Araki, Shoko, et al.
Published: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
by: Vieting, Peter, et al.
Published: (2023)
by: Vieting, Peter, et al.
Published: (2023)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
by: Li, Sirui, et al.
Published: (2025)
by: Li, Sirui, et al.
Published: (2025)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
by: von Neumann, Thilo, et al.
Published: (2025)
by: von Neumann, Thilo, et al.
Published: (2025)
Error Analysis in a Modular Meeting Transcription System
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
by: Huang, Wen-Chin, et al.
Published: (2024)
by: Huang, Wen-Chin, et al.
Published: (2024)
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
by: Meise, Adrian, et al.
Published: (2025)
by: Meise, Adrian, et al.
Published: (2025)
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
by: Boeddeker, Christoph, et al.
Published: (2024)
by: Boeddeker, Christoph, et al.
Published: (2024)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
by: Kirdey, Stanislav
Published: (2025)
by: Kirdey, Stanislav
Published: (2025)
NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
by: Ragano, Alessandro, et al.
Published: (2023)
by: Ragano, Alessandro, et al.
Published: (2023)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
by: Close, George, et al.
Published: (2024)
by: Close, George, et al.
Published: (2024)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
by: Suda, Hitoshi, et al.
Published: (2025)
by: Suda, Hitoshi, et al.
Published: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
by: Zhu, Han, et al.
Published: (2025)
by: Zhu, Han, et al.
Published: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
by: Jung, Jaemin, et al.
Published: (2024)
by: Jung, Jaemin, et al.
Published: (2024)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
by: Dai, Shuqi, et al.
Published: (2025)
by: Dai, Shuqi, et al.
Published: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
by: Huo, Mingyue, et al.
Published: (2025)
by: Huo, Mingyue, et al.
Published: (2025)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
by: Kunešová, Marie, et al.
Published: (2025)
by: Kunešová, Marie, et al.
Published: (2025)
Objective Measurements of Voice Quality
by: Dhamyal, Hira, et al.
Published: (2024)
by: Dhamyal, Hira, et al.
Published: (2024)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
by: Wang, Jiahe, et al.
Published: (2025)
by: Wang, Jiahe, et al.
Published: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
by: Zheng, Zhisheng, et al.
Published: (2025)
by: Zheng, Zhisheng, et al.
Published: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
by: Byun, Kyungguen, et al.
Published: (2025)
by: Byun, Kyungguen, et al.
Published: (2025)
Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality
by: Delgado, Pablo M., et al.
Published: (2025)
by: Delgado, Pablo M., et al.
Published: (2025)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
by: Trachu, Thanapat, et al.
Published: (2025)
by: Trachu, Thanapat, et al.
Published: (2025)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
by: Khan, Muhammad Salman, et al.
Published: (2024)
by: Khan, Muhammad Salman, et al.
Published: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
by: Li, Xuyuan, et al.
Published: (2024)
by: Li, Xuyuan, et al.
Published: (2024)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
by: Bargum, Anders R., et al.
Published: (2024)
by: Bargum, Anders R., et al.
Published: (2024)
Similar Items
-
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
by: Rautenberg, Frederik, et al.
Published: (2026) -
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
by: Rautenberg, Frederik, et al.
Published: (2025) -
Towards Frame-level Quality Predictions of Synthetic Speech
by: Kuhlmann, Michael, et al.
Published: (2025) -
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
by: Kuhlmann, Michael, et al.
Published: (2026) -
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
by: Kuhlmann, Michael, et al.
Published: (2026)