Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Liu, Tianyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
von: Kimelman, Robert G.
Veröffentlicht: (2024)
von: Kimelman, Robert G.
Veröffentlicht: (2024)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
Multimodal Contextualized Semantic Parsing from Speech
von: Voas, Jordan, et al.
Veröffentlicht: (2024)
von: Voas, Jordan, et al.
Veröffentlicht: (2024)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
von: Wang, Xihuai, et al.
Veröffentlicht: (2025)
Spontaneous Informal Speech Dataset for Punctuation Restoration
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
von: Liu, Xing Yi, et al.
Veröffentlicht: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
Are Expressions for Music Emotions the Same Across Cultures?
von: Celen, Elif, et al.
Veröffentlicht: (2025)
von: Celen, Elif, et al.
Veröffentlicht: (2025)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
von: Hui, Macarious, et al.
Veröffentlicht: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
Navigating Speech Recording Collections with AI-Generated Illustrations
von: Håland, Sirina, et al.
Veröffentlicht: (2025)
von: Håland, Sirina, et al.
Veröffentlicht: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
von: Snoubara, Abdul Aziz, et al.
Veröffentlicht: (2026)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2025)
von: Zheng, Naijun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025) -
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024) -
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025) -
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)