SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Shuang, Jiang, Yuhua, Zhou, Zineng, Liu, Dawei, Tao, Wang, Zhang, Linfeng, Qi, Biqing, Zhou, Bowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912767776129024
author Cheng, Shuang
Jiang, Yuhua
Zhou, Zineng
Liu, Dawei
Tao, Wang
Zhang, Linfeng
Qi, Biqing
Zhou, Bowen
author_facet Cheng, Shuang
Jiang, Yuhua
Zhou, Zineng
Liu, Dawei
Tao, Wang
Zhang, Linfeng
Qi, Biqing
Zhou, Bowen
contents Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been limited by high training cost, slow convergence, and instability, which have so far kept it behind strong autoregressive (AR) baselines. We present \textbf{SDAR-VL}, the first systematic application of block-wise discrete diffusion to large-scale vision-language understanding (VLU), together with an \emph{integrated framework for efficient and stable training}. This framework unifies three components: (1) \textbf{Asynchronous Block-wise Noise Scheduling} to diversify supervision within each batch; (2) \textbf{Effective Mask Ratio Scaling} for unbiased loss normalization under stochastic masking; and (3) a \textbf{Progressive Beta Noise Curriculum} that increases effective mask coverage while preserving corruption diversity. Experiments on 21 single-image, multi-image, and video benchmarks show that SDAR-VL consistently improves \emph{training efficiency}, \emph{convergence stability}, and \emph{task performance} over conventional block diffusion. On this evaluation suite, SDAR-VL sets a new state of the art among diffusion-based vision-language models and, under matched settings, matches or surpasses strong AR baselines such as LLaVA-OneVision as well as the global diffusion baseline LLaDA-V, establishing block-wise diffusion as a practical backbone for VLU.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14068
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
Cheng, Shuang
Jiang, Yuhua
Zhou, Zineng
Liu, Dawei
Tao, Wang
Zhang, Linfeng
Qi, Biqing
Zhou, Bowen
Computer Vision and Pattern Recognition
Artificial Intelligence
Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been limited by high training cost, slow convergence, and instability, which have so far kept it behind strong autoregressive (AR) baselines. We present \textbf{SDAR-VL}, the first systematic application of block-wise discrete diffusion to large-scale vision-language understanding (VLU), together with an \emph{integrated framework for efficient and stable training}. This framework unifies three components: (1) \textbf{Asynchronous Block-wise Noise Scheduling} to diversify supervision within each batch; (2) \textbf{Effective Mask Ratio Scaling} for unbiased loss normalization under stochastic masking; and (3) a \textbf{Progressive Beta Noise Curriculum} that increases effective mask coverage while preserving corruption diversity. Experiments on 21 single-image, multi-image, and video benchmarks show that SDAR-VL consistently improves \emph{training efficiency}, \emph{convergence stability}, and \emph{task performance} over conventional block diffusion. On this evaluation suite, SDAR-VL sets a new state of the art among diffusion-based vision-language models and, under matched settings, matches or surpasses strong AR baselines such as LLaVA-OneVision as well as the global diffusion baseline LLaDA-V, establishing block-wise diffusion as a practical backbone for VLU.
title SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.14068