Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Hualei, Li, Na, Wang, Chuke, Wu, Shu, Li, Zhifeng, Yu, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
di: Li, Na, et al.
Pubblicazione: (2025)
di: Li, Na, et al.
Pubblicazione: (2025)
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
di: Shen, Shengfan, et al.
Pubblicazione: (2026)
di: Shen, Shengfan, et al.
Pubblicazione: (2026)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
Zero-shot Cross-lingual Voice Transfer for TTS
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
di: Chen, Zhiyong, et al.
Pubblicazione: (2024)
di: Chen, Zhiyong, et al.
Pubblicazione: (2024)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
di: Deng, Wei, et al.
Pubblicazione: (2025)
di: Deng, Wei, et al.
Pubblicazione: (2025)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
di: Chen, Qianniu, et al.
Pubblicazione: (2025)
di: Chen, Qianniu, et al.
Pubblicazione: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
di: Chen, Dekun, et al.
Pubblicazione: (2026)
di: Chen, Dekun, et al.
Pubblicazione: (2026)
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
di: Wu, Weihao, et al.
Pubblicazione: (2025)
di: Wu, Weihao, et al.
Pubblicazione: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
di: Wang, Yuxiang, et al.
Pubblicazione: (2026)
di: Wang, Yuxiang, et al.
Pubblicazione: (2026)
IndexTTS 2.5 Technical Report
di: Li, Yunpei, et al.
Pubblicazione: (2026)
di: Li, Yunpei, et al.
Pubblicazione: (2026)
EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
di: Wu, Shu, et al.
Pubblicazione: (2025)
di: Wu, Shu, et al.
Pubblicazione: (2025)
VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
di: Krishnamurthy, Bharath, et al.
Pubblicazione: (2026)
di: Krishnamurthy, Bharath, et al.
Pubblicazione: (2026)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
di: Gan, Lu, et al.
Pubblicazione: (2025)
di: Gan, Lu, et al.
Pubblicazione: (2025)
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
di: Arora, Akshit, et al.
Pubblicazione: (2024)
di: Arora, Akshit, et al.
Pubblicazione: (2024)
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
di: Wang, Hualei, et al.
Pubblicazione: (2025)
di: Wang, Hualei, et al.
Pubblicazione: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
di: Zeldes, Ella, et al.
Pubblicazione: (2024)
di: Zeldes, Ella, et al.
Pubblicazione: (2024)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
di: Zhang, Zhedong, et al.
Pubblicazione: (2025)
di: Zhang, Zhedong, et al.
Pubblicazione: (2025)
Mobile Recording Device Recognition Based Cross-Scale and Multi-Level Representation Learning
di: Zeng, Chunyan, et al.
Pubblicazione: (2024)
di: Zeng, Chunyan, et al.
Pubblicazione: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
di: Xue, Heyang, et al.
Pubblicazione: (2025)
di: Xue, Heyang, et al.
Pubblicazione: (2025)
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
di: Chen, Junyang, et al.
Pubblicazione: (2026)
di: Chen, Junyang, et al.
Pubblicazione: (2026)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
GLM-TTS Technical Report
di: Cui, Jiayan, et al.
Pubblicazione: (2025)
di: Cui, Jiayan, et al.
Pubblicazione: (2025)
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
di: Manku, Ruskin Raj, et al.
Pubblicazione: (2025)
di: Manku, Ruskin Raj, et al.
Pubblicazione: (2025)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
VoxMind: An End-to-End Agentic Spoken Dialogue System
di: Liang, Tianle, et al.
Pubblicazione: (2026)
di: Liang, Tianle, et al.
Pubblicazione: (2026)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
Documenti analoghi
-
USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
di: Li, Na, et al.
Pubblicazione: (2025) -
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
di: Shen, Shengfan, et al.
Pubblicazione: (2026) -
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
di: Zhou, Yixuan, et al.
Pubblicazione: (2025) -
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025) -
Zero-shot Cross-lingual Voice Transfer for TTS
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)