Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhedong, Li, Liang, Yan, Chenggang, Liu, Chunshan, Hengel, Anton van den, Qi, Yuankai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
por: Cong, Gaoxiang, et al.
Publicado: (2025)
por: Cong, Gaoxiang, et al.
Publicado: (2025)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2024)
por: Cong, Gaoxiang, et al.
Publicado: (2024)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Prosody Analysis of Audiobooks
por: Pethe, Charuta, et al.
Publicado: (2023)
por: Pethe, Charuta, et al.
Publicado: (2023)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
por: Koriyama, Tomoki
Publicado: (2025)
por: Koriyama, Tomoki
Publicado: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
por: Eren, Eray, et al.
Publicado: (2025)
por: Eren, Eray, et al.
Publicado: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
por: Jiang, Yuepeng, et al.
Publicado: (2024)
por: Jiang, Yuepeng, et al.
Publicado: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025)
por: Thinakaran, Preethi, et al.
Publicado: (2025)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
por: Liu, Xiaoyang, et al.
Publicado: (2025)
por: Liu, Xiaoyang, et al.
Publicado: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
por: Qi, Tianhua, et al.
Publicado: (2024)
por: Qi, Tianhua, et al.
Publicado: (2024)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
por: Shin, Seungyoun, et al.
Publicado: (2025)
por: Shin, Seungyoun, et al.
Publicado: (2025)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
por: Borodin, Kirill, et al.
Publicado: (2026)
por: Borodin, Kirill, et al.
Publicado: (2026)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
por: Du, Chenpeng, et al.
Publicado: (2021)
por: Du, Chenpeng, et al.
Publicado: (2021)
Benchmarking Prosody Encoding in Discrete Speech Tokens
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
por: Shu, Yangyang, et al.
Publicado: (2024)
por: Shu, Yangyang, et al.
Publicado: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
Generating High-quality Symbolic Music Using Fine-grained Discriminators
por: Zhang, Zhedong, et al.
Publicado: (2024)
por: Zhang, Zhedong, et al.
Publicado: (2024)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
por: Goel, Arnav, et al.
Publicado: (2024)
por: Goel, Arnav, et al.
Publicado: (2024)
Evaluation of Virtual Acoustic Environments with Different Acoustic Level of Detail
por: Fichna, Stefan, et al.
Publicado: (2023)
por: Fichna, Stefan, et al.
Publicado: (2023)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
por: Ulgen, Ismail Rasim, et al.
Publicado: (2026)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2026)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
por: Borodin, Kirill, et al.
Publicado: (2025)
por: Borodin, Kirill, et al.
Publicado: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
por: He, Xiangheng, et al.
Publicado: (2024)
por: He, Xiangheng, et al.
Publicado: (2024)
Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
por: Yin, Jie, et al.
Publicado: (2024)
por: Yin, Jie, et al.
Publicado: (2024)
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
por: Liang, Yun, et al.
Publicado: (2024)
por: Liang, Yun, et al.
Publicado: (2024)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
por: Chevi, Rendi, et al.
Publicado: (2024)
por: Chevi, Rendi, et al.
Publicado: (2024)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
por: Lee, Kyowoon, et al.
Publicado: (2025)
por: Lee, Kyowoon, et al.
Publicado: (2025)
Usefulness of Emotional Prosody in Neural Machine Translation
por: Brazier, Charles, et al.
Publicado: (2024)
por: Brazier, Charles, et al.
Publicado: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
On Time Delay Interpolation for Improved Acoustic Reflector Localization
por: Rosseel, Hannes, et al.
Publicado: (2025)
por: Rosseel, Hannes, et al.
Publicado: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
por: Han, Wooseok, et al.
Publicado: (2024)
por: Han, Wooseok, et al.
Publicado: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
por: Zhao, Junchuan, et al.
Publicado: (2025)
por: Zhao, Junchuan, et al.
Publicado: (2025)
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
por: Hira, Medha, et al.
Publicado: (2024)
por: Hira, Medha, et al.
Publicado: (2024)
Prosody of speech production in latent post-stroke aphasia
por: Zhang, Cong, et al.
Publicado: (2024)
por: Zhang, Cong, et al.
Publicado: (2024)
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
por: Miao, Xiaoxiao, et al.
Publicado: (2024)
Cross-Domain Knowledge Transfer for Underwater Acoustic Classification Using Pre-trained Models
por: Mohammadi, Amirmohammad, et al.
Publicado: (2024)
por: Mohammadi, Amirmohammad, et al.
Publicado: (2024)
Ejemplares similares
-
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
por: Cong, Gaoxiang, et al.
Publicado: (2025) -
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2024) -
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
por: Chen, Zhengyang, et al.
Publicado: (2024) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023) -
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
por: Liu, Rui, et al.
Publicado: (2024)