Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
Fuente:
arXiv
Salvato in:
| Autori principali: | Yi, June Young, Kim, Hyeongju, Lee, Juheon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training Flow Matching Models with Reliable Labels via Self-Purification
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
WildSpoof Challenge Evaluation Plan
di: Wu, Yihan, et al.
Pubblicazione: (2025)
di: Wu, Yihan, et al.
Pubblicazione: (2025)
DFKI-Speech System for WildSpoof Challenge: A robust framework for SASV In-the-Wild
di: Das, Arnab, et al.
Pubblicazione: (2026)
di: Das, Arnab, et al.
Pubblicazione: (2026)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
di: Kim, Hyeongju, et al.
Pubblicazione: (2025)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
di: Yang, Jinhyeok, et al.
Pubblicazione: (2026)
di: Yang, Jinhyeok, et al.
Pubblicazione: (2026)
IndexTTS 2.5 Technical Report
di: Li, Yunpei, et al.
Pubblicazione: (2026)
di: Li, Yunpei, et al.
Pubblicazione: (2026)
MOSS-TTS Technical Report
di: Gong, Yitian, et al.
Pubblicazione: (2026)
di: Gong, Yitian, et al.
Pubblicazione: (2026)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
di: Luo, Dan, et al.
Pubblicazione: (2025)
di: Luo, Dan, et al.
Pubblicazione: (2025)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
A2TTS: TTS for Low Resource Indian Languages
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track
di: Sun, Renhe, et al.
Pubblicazione: (2026)
di: Sun, Renhe, et al.
Pubblicazione: (2026)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
di: Wang, Cong, et al.
Pubblicazione: (2025)
di: Wang, Cong, et al.
Pubblicazione: (2025)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
di: Han, Wooseok, et al.
Pubblicazione: (2024)
di: Han, Wooseok, et al.
Pubblicazione: (2024)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
di: Li, Bohan, et al.
Pubblicazione: (2024)
di: Li, Bohan, et al.
Pubblicazione: (2024)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
di: Chen, Qianniu, et al.
Pubblicazione: (2025)
di: Chen, Qianniu, et al.
Pubblicazione: (2025)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input
di: Liu, Changsong, et al.
Pubblicazione: (2026)
di: Liu, Changsong, et al.
Pubblicazione: (2026)
PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion
di: Pankov, Vikentii, et al.
Pubblicazione: (2026)
di: Pankov, Vikentii, et al.
Pubblicazione: (2026)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
Differentiable Reward Optimization for LLM based TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
A Non-autoregressive Model for Joint STT and TTS
di: Sunder, Vishal, et al.
Pubblicazione: (2025)
di: Sunder, Vishal, et al.
Pubblicazione: (2025)
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
di: Lee, Kyowoon, et al.
Pubblicazione: (2025)
di: Lee, Kyowoon, et al.
Pubblicazione: (2025)
TTS-1 Technical Report
di: Atamanenko, Oleg, et al.
Pubblicazione: (2025)
di: Atamanenko, Oleg, et al.
Pubblicazione: (2025)
When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
di: Purwar, Anupam, et al.
Pubblicazione: (2026)
di: Purwar, Anupam, et al.
Pubblicazione: (2026)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2025)
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
di: Shin, Seungyoun, et al.
Pubblicazione: (2025)
di: Shin, Seungyoun, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
di: Liu, Chenlin, et al.
Pubblicazione: (2025)
di: Liu, Chenlin, et al.
Pubblicazione: (2025)
SpoofCeleb: Speech Deepfake Detection and SASV In The Wild
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
di: Yang, Tianle, et al.
Pubblicazione: (2026)
di: Yang, Tianle, et al.
Pubblicazione: (2026)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
di: Liu, Hanwen, et al.
Pubblicazione: (2026)
di: Liu, Hanwen, et al.
Pubblicazione: (2026)
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS
di: Seo, Deokjin, et al.
Pubblicazione: (2026)
di: Seo, Deokjin, et al.
Pubblicazione: (2026)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
di: Qi, Xin, et al.
Pubblicazione: (2024)
di: Qi, Xin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Training Flow Matching Models with Reliable Labels via Self-Purification
di: Kim, Hyeongju, et al.
Pubblicazione: (2025) -
WildSpoof Challenge Evaluation Plan
di: Wu, Yihan, et al.
Pubblicazione: (2025) -
DFKI-Speech System for WildSpoof Challenge: A robust framework for SASV In-the-Wild
di: Das, Arnab, et al.
Pubblicazione: (2026) -
Length-Aware Rotary Position Embedding for Text-Speech Alignment
di: Kim, Hyeongju, et al.
Pubblicazione: (2025) -
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)