Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhedong, Li, Liang, Yan, Chenggang, Liu, Chunshan, Hengel, Anton van den, Qi, Yuankai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Prosody Analysis of Audiobooks
von: Pethe, Charuta, et al.
Veröffentlicht: (2023)
von: Pethe, Charuta, et al.
Veröffentlicht: (2023)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
von: Thinakaran, Preethi, et al.
Veröffentlicht: (2025)
von: Thinakaran, Preethi, et al.
Veröffentlicht: (2025)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
von: Shin, Seungyoun, et al.
Veröffentlicht: (2025)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Generating High-quality Symbolic Music Using Fine-grained Discriminators
von: Zhang, Zhedong, et al.
Veröffentlicht: (2024)
von: Zhang, Zhedong, et al.
Veröffentlicht: (2024)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
Evaluation of Virtual Acoustic Environments with Different Acoustic Level of Detail
von: Fichna, Stefan, et al.
Veröffentlicht: (2023)
von: Fichna, Stefan, et al.
Veröffentlicht: (2023)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
von: Yin, Jie, et al.
Veröffentlicht: (2024)
von: Yin, Jie, et al.
Veröffentlicht: (2024)
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
von: Liang, Yun, et al.
Veröffentlicht: (2024)
von: Liang, Yun, et al.
Veröffentlicht: (2024)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
von: Chevi, Rendi, et al.
Veröffentlicht: (2024)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
von: Lee, Kyowoon, et al.
Veröffentlicht: (2025)
von: Lee, Kyowoon, et al.
Veröffentlicht: (2025)
Usefulness of Emotional Prosody in Neural Machine Translation
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
von: Brazier, Charles, et al.
Veröffentlicht: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoxu, et al.
Veröffentlicht: (2025)
On Time Delay Interpolation for Improved Acoustic Reflector Localization
von: Rosseel, Hannes, et al.
Veröffentlicht: (2025)
von: Rosseel, Hannes, et al.
Veröffentlicht: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
von: Hira, Medha, et al.
Veröffentlicht: (2024)
von: Hira, Medha, et al.
Veröffentlicht: (2024)
Prosody of speech production in latent post-stroke aphasia
von: Zhang, Cong, et al.
Veröffentlicht: (2024)
von: Zhang, Cong, et al.
Veröffentlicht: (2024)
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
Cross-Domain Knowledge Transfer for Underwater Acoustic Classification Using Pre-trained Models
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2025) -
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024) -
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023) -
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)