Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Juncheng, Xu, Chao, Yu, Cheng, Hu, Zhe, Xie, Haoyu, Yu, Guoqi, Shang, Lei, Wang, Shujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914076616032256
author Wang, Juncheng
Xu, Chao
Yu, Cheng
Hu, Zhe
Xie, Haoyu
Yu, Guoqi
Shang, Lei
Wang, Shujun
author_facet Wang, Juncheng
Xu, Chao
Yu, Cheng
Hu, Zhe
Xie, Haoyu
Yu, Guoqi
Shang, Lei
Wang, Shujun
contents While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04577
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
Wang, Juncheng
Xu, Chao
Yu, Cheng
Hu, Zhe
Xie, Haoyu
Yu, Guoqi
Shang, Lei
Wang, Shujun
Sound
Machine Learning
Multimedia
Audio and Speech Processing
While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.
title Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
topic Sound
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2510.04577