VisualSpeech: Enhancing Prosody Modeling in TTS Using Video
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Que, Shumin, Ragni, Anton |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Emphasis Sensitivity in Speech Representations
von: Cassini, Shaun, et al.
Veröffentlicht: (2025)
von: Cassini, Shaun, et al.
Veröffentlicht: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
How I Built ASR for Endangered Languages with a Spoken Dictionary
von: Bartley, Christopher, et al.
Veröffentlicht: (2025)
von: Bartley, Christopher, et al.
Veröffentlicht: (2025)
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
von: Qian, Kaizhi, et al.
Veröffentlicht: (2025)
von: Qian, Kaizhi, et al.
Veröffentlicht: (2025)
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
von: Sasu, David, et al.
Veröffentlicht: (2025)
von: Sasu, David, et al.
Veröffentlicht: (2025)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
The Prosody of Emojis
von: Zhou, Giulio, et al.
Veröffentlicht: (2025)
von: Zhou, Giulio, et al.
Veröffentlicht: (2025)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
von: He, Jiaxu, et al.
Veröffentlicht: (2026)
von: He, Jiaxu, et al.
Veröffentlicht: (2026)
Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases
von: Zhou, Giulio, et al.
Veröffentlicht: (2024)
von: Zhou, Giulio, et al.
Veröffentlicht: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
The Role of Prosody in Spoken Question Answering
von: Chi, Jie, et al.
Veröffentlicht: (2025)
von: Chi, Jie, et al.
Veröffentlicht: (2025)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
von: Al-Sabbagh, Rania
Veröffentlicht: (2026)
von: Al-Sabbagh, Rania
Veröffentlicht: (2026)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input Adaptation
von: Chivereanu, Radu-Gabriel, et al.
Veröffentlicht: (2025)
von: Chivereanu, Radu-Gabriel, et al.
Veröffentlicht: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation
von: Jagadeeshan, Manoj Balaji, et al.
Veröffentlicht: (2026)
von: Jagadeeshan, Manoj Balaji, et al.
Veröffentlicht: (2026)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Improving French Synthetic Speech Quality via SSML Prosody Control
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
von: Ouali, Nassima Ould, et al.
Veröffentlicht: (2025)
JaiTTS: A Thai Voice Cloning Model
von: Karnjanaekarin, Jullajak, et al.
Veröffentlicht: (2026)
von: Karnjanaekarin, Jullajak, et al.
Veröffentlicht: (2026)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
Semantic Prosody in Machine Translation: the English-Chinese Case of Passive Structures
von: Ma, Xinyue, et al.
Veröffentlicht: (2025)
von: Ma, Xinyue, et al.
Veröffentlicht: (2025)
Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages
von: Pathak, Utkarsh, et al.
Veröffentlicht: (2025)
von: Pathak, Utkarsh, et al.
Veröffentlicht: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
von: Chou, Cheng-Kang, et al.
Veröffentlicht: (2025)
von: Chou, Cheng-Kang, et al.
Veröffentlicht: (2025)
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Emphasis Sensitivity in Speech Representations
von: Cassini, Shaun, et al.
Veröffentlicht: (2025) -
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025) -
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024) -
How I Built ASR for Endangered Languages with a Spoken Dictionary
von: Bartley, Christopher, et al.
Veröffentlicht: (2025) -
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
von: Qian, Kaizhi, et al.
Veröffentlicht: (2025)