Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Wenqian, Li, Xiao-Hui, Tan, Daxin, Zheng, Qiyong, King, Irwin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
von: Borodin, Kirill, et al.
Veröffentlicht: (2026)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
Recent Advances in Speech Language Models: A Survey
von: Cui, Wenqian, et al.
Veröffentlicht: (2024)
von: Cui, Wenqian, et al.
Veröffentlicht: (2024)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
von: Borodin, Kirill, et al.
Veröffentlicht: (2025)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
von: Oh, Hyung-Seok, et al.
Veröffentlicht: (2023)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
MoodLoopGP: Generating Emotion-Conditioned Loop Tablature Music with Multi-Granular Features
von: Cui, Wenqian, et al.
Veröffentlicht: (2024)
von: Cui, Wenqian, et al.
Veröffentlicht: (2024)
Leveraging LLM and Text-Queried Separation for Noise-Robust Sound Event Detection
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
SSR: Alignment-Aware Modality Connector for Speech Language Models
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
von: Cui, Wenqian, et al.
Veröffentlicht: (2025) -
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024) -
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026) -
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025) -
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)