NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Qichao, Meng, Ziqiao, Cui, Wenqian, Zhang, Yifei, Wu, Pengcheng, Wu, Bingzhe, King, Irwin, Chen, Liang, Zhao, Peilin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
di: Cui, Wenqian, et al.
Pubblicazione: (2025)
di: Cui, Wenqian, et al.
Pubblicazione: (2025)
Recent Advances in Speech Language Models: A Survey
di: Cui, Wenqian, et al.
Pubblicazione: (2024)
di: Cui, Wenqian, et al.
Pubblicazione: (2024)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
di: Cui, Wenqian, et al.
Pubblicazione: (2025)
di: Cui, Wenqian, et al.
Pubblicazione: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
di: Lu, Haitian, et al.
Pubblicazione: (2025)
di: Lu, Haitian, et al.
Pubblicazione: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
di: Qin, Chengyuan, et al.
Pubblicazione: (2025)
di: Qin, Chengyuan, et al.
Pubblicazione: (2025)
Next Tokens Denoising for Speech Synthesis
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
di: Vendrame, Katia, et al.
Pubblicazione: (2025)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
di: Ao, Junyi, et al.
Pubblicazione: (2024)
di: Ao, Junyi, et al.
Pubblicazione: (2024)
MoodLoopGP: Generating Emotion-Conditioned Loop Tablature Music with Multi-Granular Features
di: Cui, Wenqian, et al.
Pubblicazione: (2024)
di: Cui, Wenqian, et al.
Pubblicazione: (2024)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
di: Ivry, Amir, et al.
Pubblicazione: (2026)
di: Ivry, Amir, et al.
Pubblicazione: (2026)
Aligning Spoken Dialogue Models from User Interactions
di: Wu, Anne, et al.
Pubblicazione: (2025)
di: Wu, Anne, et al.
Pubblicazione: (2025)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
NAST: Noise Aware Speech Tokenization for Speech Language Models
di: Messica, Shoval, et al.
Pubblicazione: (2024)
di: Messica, Shoval, et al.
Pubblicazione: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
di: Zink, Oswald, et al.
Pubblicazione: (2024)
di: Zink, Oswald, et al.
Pubblicazione: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
di: Wu, Jingyao, et al.
Pubblicazione: (2026)
di: Wu, Jingyao, et al.
Pubblicazione: (2026)
Acoustic BPE for Speech Generation with Discrete Tokens
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
RepCodec: A Speech Representation Codec for Speech Tokenization
di: Huang, Zhichao, et al.
Pubblicazione: (2023)
di: Huang, Zhichao, et al.
Pubblicazione: (2023)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
di: Xue, Ke, et al.
Pubblicazione: (2026)
di: Xue, Ke, et al.
Pubblicazione: (2026)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
di: Cui, Zhongjian, et al.
Pubblicazione: (2025)
di: Cui, Zhongjian, et al.
Pubblicazione: (2025)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
di: Huang, Zhijie, et al.
Pubblicazione: (2026)
di: Huang, Zhijie, et al.
Pubblicazione: (2026)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
di: Wang, Dingdong, et al.
Pubblicazione: (2024)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
di: Cui, Wenqian, et al.
Pubblicazione: (2025) -
Recent Advances in Speech Language Models: A Survey
di: Cui, Wenqian, et al.
Pubblicazione: (2024) -
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
di: Cui, Wenqian, et al.
Pubblicazione: (2025) -
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025) -
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026)