VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gudmalwar, Ashishkumar, Shah, Nirmesh, Akarsh, Sai, Wasnik, Pankaj, Shah, Rajiv Ratn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?
von: Dasare, Ashwini, et al.
Veröffentlicht: (2026)
von: Dasare, Ashwini, et al.
Veröffentlicht: (2026)
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
von: Kumar, Lokesh, et al.
Veröffentlicht: (2026)
von: Kumar, Lokesh, et al.
Veröffentlicht: (2026)
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning
von: Mhaskar, Shivam Ratnakant, et al.
Veröffentlicht: (2024)
von: Mhaskar, Shivam Ratnakant, et al.
Veröffentlicht: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
von: Tripathi, Kumud, et al.
Veröffentlicht: (2025)
von: Tripathi, Kumud, et al.
Veröffentlicht: (2025)
Compact Neural TTS Voices for Accessibility
von: Jain, Kunal, et al.
Veröffentlicht: (2025)
von: Jain, Kunal, et al.
Veröffentlicht: (2025)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
Text-Driven Voice Conversion via Latent State-Space Modeling
von: Li, Wen, et al.
Veröffentlicht: (2025)
von: Li, Wen, et al.
Veröffentlicht: (2025)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
von: Wang, Chunhui, et al.
Veröffentlicht: (2024)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024) -
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024) -
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025) -
Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?
von: Dasare, Ashwini, et al.
Veröffentlicht: (2026) -
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
von: Kumar, Lokesh, et al.
Veröffentlicht: (2026)