CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Zihao, Wu, Wen, Zhang, Chao, Wu, Mengyue, Xu, Xuenan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
von: Xie, Zeyu, et al.
Veröffentlicht: (2026)
von: Xie, Zeyu, et al.
Veröffentlicht: (2026)
When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
Enhancing Audio Generation Diversity with Visual Information
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
Unified Pathological Speech Analysis with Prompt Tuning
von: Yang, Fei, et al.
Veröffentlicht: (2024)
von: Yang, Fei, et al.
Veröffentlicht: (2024)
ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal Processing Analysis for Sounds, & Strains Emerging Technology
von: Ali, Ratul, et al.
Veröffentlicht: (2023)
von: Ali, Ratul, et al.
Veröffentlicht: (2023)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025) -
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024) -
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024) -
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025) -
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)