DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lou, Yuxuan, Wu, Ziming, Wang, Yaochen, Liu, Yong, Ren, Yingxuan, Lai, Fuming, Lian, Shaobing, Tang, Jie, You, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DART: Distilling Autoregressive Reasoning to Silent Thought
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Long-Form Speech Generation with Spoken Language Models
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
von: Li, Chen, et al.
Veröffentlicht: (2025)
von: Li, Chen, et al.
Veröffentlicht: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought
von: Wang, Jianwei, et al.
Veröffentlicht: (2025)
von: Wang, Jianwei, et al.
Veröffentlicht: (2025)
A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction
von: Wang, Qing, et al.
Veröffentlicht: (2026)
von: Wang, Qing, et al.
Veröffentlicht: (2026)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Sample-Efficient Diffusion for Text-To-Speech Synthesis
von: Lovelace, Justin, et al.
Veröffentlicht: (2024)
von: Lovelace, Justin, et al.
Veröffentlicht: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
von: Yemini, Yochai, et al.
Veröffentlicht: (2023)
von: Yemini, Yochai, et al.
Veröffentlicht: (2023)
P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
von: Ren, Yong, et al.
Veröffentlicht: (2025)
von: Ren, Yong, et al.
Veröffentlicht: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
von: Li, Mohan, et al.
Veröffentlicht: (2024)
von: Li, Mohan, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
von: Xue, Jun, et al.
Veröffentlicht: (2026)
von: Xue, Jun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DART: Distilling Autoregressive Reasoning to Silent Thought
von: Jiang, Nan, et al.
Veröffentlicht: (2025) -
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026) -
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025) -
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2024) -
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)