On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
Fuente:
arXiv
Saved in:
| Main Authors: | Varshavsky-Hassid, Miri, Hirsch, Roy, Cohen, Regev, Golany, Tomer, Freedman, Daniel, Rivlin, Ehud |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Analysis of Interleaved Speech-Text Language Models
by: Maimon, Gallil, et al.
Published: (2025)
by: Maimon, Gallil, et al.
Published: (2025)
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
by: Nachmani, Eliya, et al.
Published: (2023)
by: Nachmani, Eliya, et al.
Published: (2023)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
by: Shor, Joel, et al.
Published: (2023)
by: Shor, Joel, et al.
Published: (2023)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
by: Xin, Detai, et al.
Published: (2026)
by: Xin, Detai, et al.
Published: (2026)
Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
by: Ma, Zhengrui, et al.
Published: (2025)
by: Ma, Zhengrui, et al.
Published: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
by: Mousavi, Pooneh, et al.
Published: (2025)
by: Mousavi, Pooneh, et al.
Published: (2025)
Continuous Speech Tokenizer in Text To Speech
by: Li, Yixing, et al.
Published: (2024)
by: Li, Yixing, et al.
Published: (2024)
SpeechTaxi: On Multilingual Semantic Speech Classification
by: Keller, Lennart, et al.
Published: (2024)
by: Keller, Lennart, et al.
Published: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
by: Kong, Jungil, et al.
Published: (2023)
by: Kong, Jungil, et al.
Published: (2023)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
by: Futami, Hayato, et al.
Published: (2025)
by: Futami, Hayato, et al.
Published: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
by: Oh, Hyung-Seok, et al.
Published: (2023)
by: Oh, Hyung-Seok, et al.
Published: (2023)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
by: Lou, Haowei, et al.
Published: (2024)
by: Lou, Haowei, et al.
Published: (2024)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
by: Shi, Hao, et al.
Published: (2023)
by: Shi, Hao, et al.
Published: (2023)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
by: Zhao, Shengkui, et al.
Published: (2025)
by: Zhao, Shengkui, et al.
Published: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
BoSS: Beyond-Semantic Speech
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Semantic MIMO Systems for Speech-to-Text Transmission
by: Weng, Zhenzi, et al.
Published: (2024)
by: Weng, Zhenzi, et al.
Published: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
by: Zhang, Shaozuo, et al.
Published: (2025)
by: Zhang, Shaozuo, et al.
Published: (2025)
Zero-Shot Text-to-Speech for Vietnamese
by: Vu, Thi, et al.
Published: (2025)
by: Vu, Thi, et al.
Published: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
by: Liu, Henglyu, et al.
Published: (2025)
by: Liu, Henglyu, et al.
Published: (2025)
Textually Pretrained Speech Language Models
by: Hassid, Michael, et al.
Published: (2023)
by: Hassid, Michael, et al.
Published: (2023)
Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
by: Sakpiboonchit, Siratish
Published: (2025)
by: Sakpiboonchit, Siratish
Published: (2025)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
by: Li, Zhu, et al.
Published: (2025)
by: Li, Zhu, et al.
Published: (2025)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
Computational Narrative Understanding for Expressive Text-to-Speech
by: Michel, Gaspard, et al.
Published: (2025)
by: Michel, Gaspard, et al.
Published: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
by: Vendrame, Katia, et al.
Published: (2025)
by: Vendrame, Katia, et al.
Published: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
by: Kim, Heeseung, et al.
Published: (2024)
by: Kim, Heeseung, et al.
Published: (2024)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
by: Kang, Wonjune, et al.
Published: (2025)
by: Kang, Wonjune, et al.
Published: (2025)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
by: Lam, Tsz Kin, et al.
Published: (2025)
by: Lam, Tsz Kin, et al.
Published: (2025)
End-to-End Speech-to-Text Translation: A Survey
by: Sethiya, Nivedita, et al.
Published: (2023)
by: Sethiya, Nivedita, et al.
Published: (2023)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
MunTTS: A Text-to-Speech System for Mundari
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
by: Fucci, Dennis, et al.
Published: (2024)
by: Fucci, Dennis, et al.
Published: (2024)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
by: Bott, Thomas, et al.
Published: (2024)
by: Bott, Thomas, et al.
Published: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
by: Do, Cong-Thanh, et al.
Published: (2024)
by: Do, Cong-Thanh, et al.
Published: (2024)
Similar Items
-
Scaling Analysis of Interleaved Speech-Text Language Models
by: Maimon, Gallil, et al.
Published: (2025) -
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
by: Nachmani, Eliya, et al.
Published: (2023) -
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024) -
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
by: Shor, Joel, et al.
Published: (2023) -
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
by: Xin, Detai, et al.
Published: (2026)