From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Tianqiao, Li, Xueyi, Wang, Hao, Li, Haoxuan, Chen, Zhichao, Luo, Weiqi, Liu, Zitao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917360369139712
author Liu, Tianqiao
Li, Xueyi
Wang, Hao
Li, Haoxuan
Chen, Zhichao
Luo, Weiqi
Liu, Zitao
author_facet Liu, Tianqiao
Li, Xueyi
Wang, Hao
Li, Haoxuan
Chen, Zhichao
Luo, Weiqi
Liu, Zitao
contents Recent advances in large language models (LLMs) have attracted significant interest in extending their capabilities to multimodal scenarios, particularly for speech-to-speech conversational systems. However, existing multimodal models handling interleaved audio and text rely on autoregressive (AR) methods, overlooking that text depends on target-target relations whereas audio depends mainly on source-target relations. In this work, we propose Text-to-Talk (TtT), a unified audio-text framework that integrates AR text generation with non-autoregressive (NAR) audio diffusion in a single Transformer. By leveraging the any-order AR property of absorbing discrete diffusion, our approach provides a unified training objective for text and audio. To support this hybrid generation paradigm, we design a modality-aware attention mechanism that enforces causal decoding for text while allowing bidirectional modeling within audio spans, and further introduce three training strategies that reduce train-test discrepancies. During inference, TtT employs block-wise diffusion to synthesize audio in parallel while flexibly handling variable-length outputs. Comprehensive experiments on Audio-QA, ASR, AAC and speech-to-speech benchmarks show that TtT consistently surpasses strong AR and NAR baselines, with additional ablation and training-strategy analyses confirming the contribution of each component. We will open-source our models, data and code to facilitate future research in this direction.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
Liu, Tianqiao
Li, Xueyi
Wang, Hao
Li, Haoxuan
Chen, Zhichao
Luo, Weiqi
Liu, Zitao
Computation and Language
Recent advances in large language models (LLMs) have attracted significant interest in extending their capabilities to multimodal scenarios, particularly for speech-to-speech conversational systems. However, existing multimodal models handling interleaved audio and text rely on autoregressive (AR) methods, overlooking that text depends on target-target relations whereas audio depends mainly on source-target relations. In this work, we propose Text-to-Talk (TtT), a unified audio-text framework that integrates AR text generation with non-autoregressive (NAR) audio diffusion in a single Transformer. By leveraging the any-order AR property of absorbing discrete diffusion, our approach provides a unified training objective for text and audio. To support this hybrid generation paradigm, we design a modality-aware attention mechanism that enforces causal decoding for text while allowing bidirectional modeling within audio spans, and further introduce three training strategies that reduce train-test discrepancies. During inference, TtT employs block-wise diffusion to synthesize audio in parallel while flexibly handling variable-length outputs. Comprehensive experiments on Audio-QA, ASR, AAC and speech-to-speech benchmarks show that TtT consistently surpasses strong AR and NAR baselines, with additional ablation and training-strategy analyses confirming the contribution of each component. We will open-source our models, data and code to facilitate future research in this direction.
title From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training
topic Computation and Language
url https://arxiv.org/abs/2509.20072