DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuanyuan, Yang, Dongchao, Shao, Yiwen, Chen, Hangting, Zhao, Jiankun, Wu, Zhiyong, Meng, Helen, Wu, Xixin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
UniSep: Universal Target Audio Separation with Language Models at Scale
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
DENSE: Dynamic Embedding Causal Target Speech Extraction
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
von: Zhao, He, et al.
Veröffentlicht: (2024)
von: Zhao, He, et al.
Veröffentlicht: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026) -
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
von: Guo, Haohan, et al.
Veröffentlicht: (2024) -
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024) -
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)