Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Xinlu, Ray, Swayambhu Nath, Mallidi, Harish, Huang, Jia-Hong, Bellur, Ashwin, Chandak, Chander, Maruf, M., Ravichandran, Venkatesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
por: Jeon, Yejin, et al.
Publicado: (2024)
por: Jeon, Yejin, et al.
Publicado: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
por: Tankasala, Srinath, et al.
Publicado: (2023)
por: Tankasala, Srinath, et al.
Publicado: (2023)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
por: Kunešová, Marie, et al.
Publicado: (2025)
por: Kunešová, Marie, et al.
Publicado: (2025)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
por: Pankov, Vikentii, et al.
Publicado: (2023)
por: Pankov, Vikentii, et al.
Publicado: (2023)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
por: Nie, Sihang, et al.
Publicado: (2025)
por: Nie, Sihang, et al.
Publicado: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
por: Liu, Huadai, et al.
Publicado: (2023)
por: Liu, Huadai, et al.
Publicado: (2023)
SponTTS: modeling and transferring spontaneous style for TTS
por: Li, Hanzhao, et al.
Publicado: (2023)
por: Li, Hanzhao, et al.
Publicado: (2023)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
por: Ko, Myeongjin, et al.
Publicado: (2023)
por: Ko, Myeongjin, et al.
Publicado: (2023)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
por: Meng, Ming, et al.
Publicado: (2025)
por: Meng, Ming, et al.
Publicado: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
por: Liu, Zhijun, et al.
Publicado: (2024)
por: Liu, Zhijun, et al.
Publicado: (2024)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
por: Wang, Chunhui, et al.
Publicado: (2024)
por: Wang, Chunhui, et al.
Publicado: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
Diffusion-Based Adversarial Purification for Speaker Verification
por: Bai, Yibo, et al.
Publicado: (2023)
por: Bai, Yibo, et al.
Publicado: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
por: Truong, Duc-Tuan, et al.
Publicado: (2023)
por: Truong, Duc-Tuan, et al.
Publicado: (2023)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
por: Han, Wooseok, et al.
Publicado: (2024)
por: Han, Wooseok, et al.
Publicado: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
por: Chen, Zhiyong, et al.
Publicado: (2024)
por: Chen, Zhiyong, et al.
Publicado: (2024)
Speaker Contrastive Learning for Source Speaker Tracing
por: Wang, Qing, et al.
Publicado: (2024)
por: Wang, Qing, et al.
Publicado: (2024)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
por: Li, Xin, et al.
Publicado: (2025)
por: Li, Xin, et al.
Publicado: (2025)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
por: Sundar, Anirudh S., et al.
Publicado: (2023)
por: Sundar, Anirudh S., et al.
Publicado: (2023)
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
por: Lin, Chaohao, et al.
Publicado: (2025)
por: Lin, Chaohao, et al.
Publicado: (2025)
Accent-VITS:accent transfer for end-to-end TTS
por: Ma, Linhan, et al.
Publicado: (2023)
por: Ma, Linhan, et al.
Publicado: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024)
por: Jung, Sunghee, et al.
Publicado: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
por: Biadsy, Fadi, et al.
Publicado: (2024)
por: Biadsy, Fadi, et al.
Publicado: (2024)
A Neural Score Follower for Computer Accompaniment of Polyphonic Musical Instruments
por: Pillay, Ashwin
Publicado: (2025)
por: Pillay, Ashwin
Publicado: (2025)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
por: Zhou, Zhenyu, et al.
Publicado: (2024)
por: Zhou, Zhenyu, et al.
Publicado: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
por: Zhang, Ke, et al.
Publicado: (2024)
por: Zhang, Ke, et al.
Publicado: (2024)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
por: You, Zhenghai, et al.
Publicado: (2025)
por: You, Zhenghai, et al.
Publicado: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
por: Wu, Shu, et al.
Publicado: (2025)
por: Wu, Shu, et al.
Publicado: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Ejemplares similares
-
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
por: Jeon, Yejin, et al.
Publicado: (2024) -
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024) -
Cross-utterance ASR Rescoring with Graph-based Label Propagation
por: Tankasala, Srinath, et al.
Publicado: (2023) -
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
por: Kunešová, Marie, et al.
Publicado: (2025) -
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
por: Wang, Tianrui, et al.
Publicado: (2025)