WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yifu, Ji, Shengpeng, Chen, Qian, Liang, Tianle, Li, Yangzhuo, Wang, Ziqing, Wang, Wen, Lu, Jingyu, Wang, Haoxiao, Pu, Xueyi, Zhuo, Fan, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
von: Liang, Tianle, et al.
Veröffentlicht: (2026)
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026)
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2026)
von: Chen, Yifu, et al.
Veröffentlicht: (2026)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
WavChat: A Survey of Spoken Dialogue Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
von: Cheng, Xize, et al.
Veröffentlicht: (2025)
EmoNews: A Spoken Dialogue System for Expressive News Conversations
von: Matsuura, Ryuki, et al.
Veröffentlicht: (2025)
von: Matsuura, Ryuki, et al.
Veröffentlicht: (2025)
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025)
von: Wu, Anne, et al.
Veröffentlicht: (2025)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
von: Fan, Wei, et al.
Veröffentlicht: (2025)
von: Fan, Wei, et al.
Veröffentlicht: (2025)
AdaFSNet: Time Series Classification Based on Convolutional Network with a Adaptive and Effective Kernel Size Configuration
von: Wang, Haoxiao, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiao, et al.
Veröffentlicht: (2024)
Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
von: Qian, Livia, et al.
Veröffentlicht: (2024)
von: Qian, Livia, et al.
Veröffentlicht: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
von: Ao, Junyi, et al.
Veröffentlicht: (2024)
von: Ao, Junyi, et al.
Veröffentlicht: (2024)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
CombAlign: Enhancing Model Expressiveness in Unsupervised Graph Alignment
von: Chen, Songyang, et al.
Veröffentlicht: (2024)
von: Chen, Songyang, et al.
Veröffentlicht: (2024)
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
von: Wang, Weijie, et al.
Veröffentlicht: (2025)
MOSS-TTSD: Text to Spoken Dialogue Generation
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
Are LLMs Robust for Spoken Dialogues?
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
Diffusion Model as a Generalist Segmentation Learner
von: Wang, Haoxiao, et al.
Veröffentlicht: (2026)
von: Wang, Haoxiao, et al.
Veröffentlicht: (2026)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
WavFlow: Audio Generation in Waveform Space
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
von: Wu, Yihao, et al.
Veröffentlicht: (2025)
TiCo: Time-Controllable Spoken Dialogue Model
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Characterizing Disease Modifying Effects Using Overall Treatment Effect Across all Post‐Baseline Visits
von: John O'Gorman, et al.
Veröffentlicht: (2024)
von: John O'Gorman, et al.
Veröffentlicht: (2024)
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
von: Xu, Jiacheng, et al.
Veröffentlicht: (2026)
von: Xu, Jiacheng, et al.
Veröffentlicht: (2026)
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
von: Nakazawa, Kazushi
Veröffentlicht: (2026)
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
von: Lee, Jonggeun, et al.
Veröffentlicht: (2026)
von: Lee, Jonggeun, et al.
Veröffentlicht: (2026)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Can Language Models Trained on Written Monologue Learn to Predict Spoken Dialogue?
von: Muhammad Umair, et al.
Veröffentlicht: (2024)
von: Muhammad Umair, et al.
Veröffentlicht: (2024)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
PTQ4SAM: Post-Training Quantization for Segment Anything
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
von: Lv, Chengtao, et al.
Veröffentlicht: (2024)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
von: Yoon, Jaeseok, et al.
Veröffentlicht: (2023)
von: Yoon, Jaeseok, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025) -
VoxMind: An End-to-End Agentic Spoken Dialogue System
von: Liang, Tianle, et al.
Veröffentlicht: (2026) -
WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
von: Li, Yangzhuo, et al.
Veröffentlicht: (2026) -
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2026) -
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)