Soundwave: Less is More for Speech-Text Alignment in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yuhao, Liu, Zhiheng, Bu, Fan, Zhang, Ruiyu, Wang, Benyou, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
von: Bu, Fan, et al.
Veröffentlicht: (2024)
von: Bu, Fan, et al.
Veröffentlicht: (2024)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
von: Lin, Liang, et al.
Veröffentlicht: (2026)
von: Lin, Liang, et al.
Veröffentlicht: (2026)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
von: Gao, Yan, et al.
Veröffentlicht: (2025)
von: Gao, Yan, et al.
Veröffentlicht: (2025)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
von: Zhou, Li, et al.
Veröffentlicht: (2026)
von: Zhou, Li, et al.
Veröffentlicht: (2026)
Cross-Attention is Half Explanation in Speech-to-Text Models
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
von: Wang, Peidong
Veröffentlicht: (2026)
von: Wang, Peidong
Veröffentlicht: (2026)
EvA: An Evidence-First Audio Understanding Paradigm for LALMs
von: Xie, Xinyuan, et al.
Veröffentlicht: (2026)
von: Xie, Xinyuan, et al.
Veröffentlicht: (2026)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
von: Han, Yichen, et al.
Veröffentlicht: (2025)
von: Han, Yichen, et al.
Veröffentlicht: (2025)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
von: Kotoge, Rikuto, et al.
Veröffentlicht: (2025)
von: Kotoge, Rikuto, et al.
Veröffentlicht: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
von: Papi, Sara, et al.
Veröffentlicht: (2026)
von: Papi, Sara, et al.
Veröffentlicht: (2026)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
MOSS-TTSD: Text to Spoken Dialogue Generation
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
von: Wang, Peng, et al.
Veröffentlicht: (2026)
von: Wang, Peng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025) -
Roadmap towards Superhuman Speech Understanding using Large Language Models
von: Bu, Fan, et al.
Veröffentlicht: (2024) -
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025) -
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025) -
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)