PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hong, Changi, Song, Yoonah, Park, Hwayoung, Bang, Chaewoon, Ku, Dayeon, Lee, Do Hyun, Kim, Hong Kook |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
par: Lee, Do Hyun, et autres
Publié: (2024)
par: Lee, Do Hyun, et autres
Publié: (2024)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
par: Sahipjohn, Neha, et autres
Publié: (2024)
par: Sahipjohn, Neha, et autres
Publié: (2024)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
par: Park, Hyun Joon, et autres
Publié: (2024)
par: Park, Hyun Joon, et autres
Publié: (2024)
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
par: Kim, Semin, et autres
Publié: (2026)
par: Kim, Semin, et autres
Publié: (2026)
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
par: Park, Hyun Joon, et autres
Publié: (2025)
par: Park, Hyun Joon, et autres
Publié: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
par: Choi, Jeongsoo, et autres
Publié: (2025)
par: Choi, Jeongsoo, et autres
Publié: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
par: Liu, Huadai, et autres
Publié: (2023)
par: Liu, Huadai, et autres
Publié: (2023)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
par: Sung-Bin, Kim, et autres
Publié: (2025)
par: Sung-Bin, Kim, et autres
Publié: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
par: Dai, Ziqi, et autres
Publié: (2025)
par: Dai, Ziqi, et autres
Publié: (2025)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
par: Lee, Yoonhyung, et autres
Publié: (2026)
par: Lee, Yoonhyung, et autres
Publié: (2026)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
par: Ma, Linhan, et autres
Publié: (2024)
par: Ma, Linhan, et autres
Publié: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
par: Abilbekov, Adal, et autres
Publié: (2024)
par: Abilbekov, Adal, et autres
Publié: (2024)
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
par: Li, Zirui, et autres
Publié: (2025)
par: Li, Zirui, et autres
Publié: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
par: Zhou, Kun, et autres
Publié: (2024)
par: Zhou, Kun, et autres
Publié: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
par: Kim, Jaehyeon, et autres
Publié: (2024)
par: Kim, Jaehyeon, et autres
Publié: (2024)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
par: Li, Zirui, et autres
Publié: (2025)
par: Li, Zirui, et autres
Publié: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
par: Guo, Hao-Han, et autres
Publié: (2025)
par: Guo, Hao-Han, et autres
Publié: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
par: Guo, Hao-Han, et autres
Publié: (2024)
par: Guo, Hao-Han, et autres
Publié: (2024)
MathReader : Text-to-Speech for Mathematical Documents
par: Hyeon, Sieun, et autres
Publié: (2025)
par: Hyeon, Sieun, et autres
Publié: (2025)
Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
par: Park, Jongyeon, et autres
Publié: (2025)
par: Park, Jongyeon, et autres
Publié: (2025)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
par: Kim, Hyeongju, et autres
Publié: (2025)
par: Kim, Hyeongju, et autres
Publié: (2025)
MunTTS: A Text-to-Speech System for Mundari
par: Gumma, Varun, et autres
Publié: (2024)
par: Gumma, Varun, et autres
Publié: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
par: Gudmalwar, Ashishkumar, et autres
Publié: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
par: Fu, Ruibo, et autres
Publié: (2024)
par: Fu, Ruibo, et autres
Publié: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
par: Guo, Yinlin, et autres
Publié: (2024)
par: Guo, Yinlin, et autres
Publié: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
par: Singh, Jaskaran, et autres
Publié: (2025)
par: Singh, Jaskaran, et autres
Publié: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
par: Park, Nohil, et autres
Publié: (2024)
par: Park, Nohil, et autres
Publié: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
par: Han, Seungu, et autres
Publié: (2026)
par: Han, Seungu, et autres
Publié: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
par: Han, Wooseok, et autres
Publié: (2024)
par: Han, Wooseok, et autres
Publié: (2024)
Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
par: Azzouz, Sofiane, et autres
Publié: (2026)
par: Azzouz, Sofiane, et autres
Publié: (2026)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
par: Xue, Heyang, et autres
Publié: (2025)
par: Xue, Heyang, et autres
Publié: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
par: Bondaruk, Łukasz, et autres
Publié: (2025)
par: Bondaruk, Łukasz, et autres
Publié: (2025)
Speech Codec Probing from Semantic and Phonetic Perspectives
par: Shi, Xuan, et autres
Publié: (2026)
par: Shi, Xuan, et autres
Publié: (2026)
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
par: Son, Sang Won, et autres
Publié: (2024)
par: Son, Sang Won, et autres
Publié: (2024)
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
par: Ku, Pin-Jui, et autres
Publié: (2025)
par: Ku, Pin-Jui, et autres
Publié: (2025)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
par: Lou, Haowei, et autres
Publié: (2025)
par: Lou, Haowei, et autres
Publié: (2025)
Documents similaires
-
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
par: Lee, Do Hyun, et autres
Publié: (2024) -
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
par: Sahipjohn, Neha, et autres
Publié: (2024) -
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
par: Park, Hyun Joon, et autres
Publié: (2024) -
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
par: Kim, Semin, et autres
Publié: (2026) -
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
par: Park, Hyun Joon, et autres
Publié: (2025)