WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yifu, Ji, Shengpeng, Chen, Qian, Liang, Tianle, Li, Yangzhuo, Wang, Ziqing, Wang, Wen, Lu, Jingyu, Wang, Haoxiao, Pu, Xueyi, Zhuo, Fan, Zhao, Zhou
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917413719638016
author Chen, Yifu
Ji, Shengpeng
Chen, Qian
Liang, Tianle
Li, Yangzhuo
Wang, Ziqing
Wang, Wen
Lu, Jingyu
Wang, Haoxiao
Pu, Xueyi
Zhuo, Fan
Zhao, Zhou
author_facet Chen, Yifu
Ji, Shengpeng
Chen, Qian
Liang, Tianle
Li, Yangzhuo
Wang, Ziqing
Wang, Wen
Lu, Jingyu
Wang, Haoxiao
Pu, Xueyi
Zhuo, Fan
Zhao, Zhou
contents End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems. However, the intelligence and expressiveness of current open-source spoken dialogue models often remain below expectations. Motivated by the success of online reinforcement learning(RL) in other domains, one might attempt to directly apply preference optimization to spoken dialogue models, yet this transfer is non-trivial. We analyze these obstacles from the perspectives of reward modeling and rollout sampling, focusing on how sparse preference supervision interacts with dense speech generation under shared-parameter updates. Based on the analysis, we propose a modality-aware adaptive post-training recipe that makes RL practical for spoken dialogue: it constrains preference updates to the semantic channel and improves acoustic behavior via explicit anchoring, while dynamically regulating their mixture from rollout statistics to avoid unreliable preference gradients. We evaluate the method across multiple spoken dialogue benchmarks and representative architectures, and observe consistent improvements in semantic quality and speech expressiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14932
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
Chen, Yifu
Ji, Shengpeng
Chen, Qian
Liang, Tianle
Li, Yangzhuo
Wang, Ziqing
Wang, Wen
Lu, Jingyu
Wang, Haoxiao
Pu, Xueyi
Zhuo, Fan
Zhao, Zhou
Artificial Intelligence
End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems. However, the intelligence and expressiveness of current open-source spoken dialogue models often remain below expectations. Motivated by the success of online reinforcement learning(RL) in other domains, one might attempt to directly apply preference optimization to spoken dialogue models, yet this transfer is non-trivial. We analyze these obstacles from the perspectives of reward modeling and rollout sampling, focusing on how sparse preference supervision interacts with dense speech generation under shared-parameter updates. Based on the analysis, we propose a modality-aware adaptive post-training recipe that makes RL practical for spoken dialogue: it constrains preference updates to the semantic channel and improves acoustic behavior via explicit anchoring, while dynamically regulating their mixture from rollout statistics to avoid unreliable preference gradients. We evaluate the method across multiple spoken dialogue benchmarks and representative architectures, and observe consistent improvements in semantic quality and speech expressiveness.
title WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
topic Artificial Intelligence
url https://arxiv.org/abs/2604.14932