Improving French Synthetic Speech Quality via SSML Prosody Control
Fuente:
arXiv
Saved in:
| Main Authors: | Ouali, Nassima Ould, Sani, Awais Hussain, Bueno, Ruben, Dauvet, Jonah, Horstmann, Tim Luka, Moulines, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
by: Geleta, Margarita, et al.
Published: (2025)
by: Geleta, Margarita, et al.
Published: (2025)
Online Correction of Dispersion Error in 2D Waveguide Meshes
by: Fontana, Federico, et al.
Published: (2000)
by: Fontana, Federico, et al.
Published: (2000)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
by: Zhu, Di, et al.
Published: (2026)
by: Zhu, Di, et al.
Published: (2026)
Dichotic harmony for the musical practice
by: Madgazin, Vadim R.
Published: (2010)
by: Madgazin, Vadim R.
Published: (2010)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
by: Wang, Shaowen, et al.
Published: (2025)
by: Wang, Shaowen, et al.
Published: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
by: Li, Pengcheng, et al.
Published: (2024)
by: Li, Pengcheng, et al.
Published: (2024)
Masked Contrastive Pre-Training Improves Music Audio Key Detection
by: Yonay, Ori, et al.
Published: (2026)
by: Yonay, Ori, et al.
Published: (2026)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
by: Mehdi, Naqcho Ali, et al.
Published: (2026)
by: Mehdi, Naqcho Ali, et al.
Published: (2026)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
by: Marmoret, Axel
Published: (2026)
by: Marmoret, Axel
Published: (2026)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
by: Kim, Minu, et al.
Published: (2025)
by: Kim, Minu, et al.
Published: (2025)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
by: Bhattacharjee, Aditya, et al.
Published: (2025)
by: Bhattacharjee, Aditya, et al.
Published: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
by: Donepudi, Dharma Teja
Published: (2025)
by: Donepudi, Dharma Teja
Published: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
by: Chen, Kuan-Yu, et al.
Published: (2025)
by: Chen, Kuan-Yu, et al.
Published: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
by: Viswanathan, Janaki, et al.
Published: (2025)
by: Viswanathan, Janaki, et al.
Published: (2025)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
by: Park, Jisang, et al.
Published: (2024)
by: Park, Jisang, et al.
Published: (2024)
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Signal-Theoretic Characterization of Waveguide Mesh Geometries for Models of Two--Dimensional Wave Propagation in Elastic Media
by: Fontana, Federico, et al.
Published: (2001)
by: Fontana, Federico, et al.
Published: (2001)
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
by: Qian, Lekai, et al.
Published: (2026)
by: Qian, Lekai, et al.
Published: (2026)
The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
by: Chen, Kuan-Yu, et al.
Published: (2026)
by: Chen, Kuan-Yu, et al.
Published: (2026)
If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
by: Nunes, Caio, et al.
Published: (2025)
by: Nunes, Caio, et al.
Published: (2025)
Synthetic Media in Multilingual MOOCs: Deepfake Tutors, Pedagogical Effects, and Ethical-Policy Challenges
by: Gazis, Alexandros, et al.
Published: (2026)
by: Gazis, Alexandros, et al.
Published: (2026)
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
by: Rekimoto, Jun, et al.
Published: (2026)
by: Rekimoto, Jun, et al.
Published: (2026)
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
by: Dhiman, Jai
Published: (2026)
by: Dhiman, Jai
Published: (2026)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
by: Shu, Hongzhi, et al.
Published: (2024)
by: Shu, Hongzhi, et al.
Published: (2024)
Generation of Musical Timbres using a Text-Guided Diffusion Model
by: Yuan, Weixuan, et al.
Published: (2025)
by: Yuan, Weixuan, et al.
Published: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
by: Mahfi, Muntahi Safwan, et al.
Published: (2025)
by: Mahfi, Muntahi Safwan, et al.
Published: (2025)
A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation
by: He, Qiqi, et al.
Published: (2026)
by: He, Qiqi, et al.
Published: (2026)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
by: Chopra, Anuradha, et al.
Published: (2025)
by: Chopra, Anuradha, et al.
Published: (2025)
Quantum-Enhanced Analysis and Grading of Vocal Performance
by: Agarwal, Rohan
Published: (2025)
by: Agarwal, Rohan
Published: (2025)
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
by: Maben, Leander Melroy, et al.
Published: (2025)
by: Maben, Leander Melroy, et al.
Published: (2025)
Can pre-trained Deep Learning models predict groove ratings?
by: Marmoret, Axel, et al.
Published: (2026)
by: Marmoret, Axel, et al.
Published: (2026)
Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
by: Samsurya, Bilkent
Published: (2025)
by: Samsurya, Bilkent
Published: (2025)
Refining music sample identification with a self-supervised graph neural network
by: Bhattacharjee, Aditya, et al.
Published: (2025)
by: Bhattacharjee, Aditya, et al.
Published: (2025)
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
by: Cauzinille, Jules, et al.
Published: (2025)
by: Cauzinille, Jules, et al.
Published: (2025)
Neural Proxies for Sound Synthesizers: Learning Perceptually Informed Preset Representations
by: Combes, Paolo, et al.
Published: (2025)
by: Combes, Paolo, et al.
Published: (2025)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
by: Cho, Hyunsung, et al.
Published: (2024)
by: Cho, Hyunsung, et al.
Published: (2024)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
by: Woodard, Brandon, et al.
Published: (2025)
by: Woodard, Brandon, et al.
Published: (2025)
REMAST: Real-time Emotion-based Music Arrangement with Soft Transition
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
GraFPrint: A GNN-Based Approach for Audio Identification
by: Bhattacharjee, Aditya, et al.
Published: (2024)
by: Bhattacharjee, Aditya, et al.
Published: (2024)
Similar Items
-
Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
by: Geleta, Margarita, et al.
Published: (2025) -
Online Correction of Dispersion Error in 2D Waveguide Meshes
by: Fontana, Federico, et al.
Published: (2000) -
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
by: Zhu, Di, et al.
Published: (2026) -
Dichotic harmony for the musical practice
by: Madgazin, Vadim R.
Published: (2010) -
Self-Improvement for Audio Large Language Model using Unlabeled Speech
by: Wang, Shaowen, et al.
Published: (2025)