DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Ye-Xin, Gu, Yu, Wei, Kun, Du, Hui-Peng, Ai, Yang, Ling, Zhen-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)
von: Liu, Fei, et al.
Veröffentlicht: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
Universal Preference-Score-based Pairwise Speech Quality Assessment
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2025)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
von: Wang, Qian, et al.
Veröffentlicht: (2024)
von: Wang, Qian, et al.
Veröffentlicht: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Rui-Chen, et al.
Veröffentlicht: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Serenade: A Singing Style Conversion Framework Based On Audio Infilling
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
von: Xie, Kun, et al.
Veröffentlicht: (2025)
von: Xie, Kun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025) -
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024) -
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024) -
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023) -
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)