PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Changi, Song, Yoonah, Park, Hwayoung, Bang, Chaewoon, Ku, Dayeon, Lee, Do Hyun, Kim, Hong Kook |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
von: Kim, Semin, et al.
Veröffentlicht: (2026)
von: Kim, Semin, et al.
Veröffentlicht: (2026)
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
von: Park, Hyun Joon, et al.
Veröffentlicht: (2025)
von: Park, Hyun Joon, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
von: Li, Zirui, et al.
Veröffentlicht: (2025)
von: Li, Zirui, et al.
Veröffentlicht: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
MathReader : Text-to-Speech for Mathematical Documents
von: Hyeon, Sieun, et al.
Veröffentlicht: (2025)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2025)
Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
von: Park, Jongyeon, et al.
Veröffentlicht: (2025)
von: Park, Jongyeon, et al.
Veröffentlicht: (2025)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
von: Singh, Jaskaran, et al.
Veröffentlicht: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
Speech Codec Probing from Semantic and Phonetic Perspectives
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
von: Son, Sang Won, et al.
Veröffentlicht: (2024)
von: Son, Sang Won, et al.
Veröffentlicht: (2024)
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2025)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2025)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024) -
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024) -
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024) -
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
von: Kim, Semin, et al.
Veröffentlicht: (2026) -
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
von: Park, Hyun Joon, et al.
Veröffentlicht: (2025)