DIFFA: Large Language Diffusion Models Can Listen and Understand

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Jiaming, Chen, Hongjie, Zhao, Shiwan, Kang, Jian, Li, Jie, Wang, Enzhi, Guo, Yujie, Sun, Haoqin, Wang, Hui, Kong, Aobo, Qin, Yong, Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912697158729728
author Zhou, Jiaming
Chen, Hongjie
Zhao, Shiwan
Kang, Jian
Li, Jie
Wang, Enzhi
Guo, Yujie
Sun, Haoqin
Wang, Hui
Kong, Aobo
Qin, Yong
Li, Xuelong
author_facet Zhou, Jiaming
Chen, Hongjie
Zhao, Shiwan
Kang, Jian
Li, Jie
Wang, Enzhi
Guo, Yujie
Sun, Haoqin
Wang, Hui
Kong, Aobo
Qin, Yong
Li, Xuelong
contents Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce \textbf{DIFFA}, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of diffusion-based language models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Our code will be available at https://github.com/NKU-HLT/DIFFA.git.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DIFFA: Large Language Diffusion Models Can Listen and Understand
Zhou, Jiaming
Chen, Hongjie
Zhao, Shiwan
Kang, Jian
Li, Jie
Wang, Enzhi
Guo, Yujie
Sun, Haoqin
Wang, Hui
Kong, Aobo
Qin, Yong
Li, Xuelong
Sound
Audio and Speech Processing
Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce \textbf{DIFFA}, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of diffusion-based language models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Our code will be available at https://github.com/NKU-HLT/DIFFA.git.
title DIFFA: Large Language Diffusion Models Can Listen and Understand
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.18452