Saved in:
Bibliographic Details
Main Authors: Zuo, Haomin, Li, Yidi, Yang, Luoxiao, Zhang, Xiaofeng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.11005
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908957926227968
author Zuo, Haomin
Li, Yidi
Yang, Luoxiao
Zhang, Xiaofeng
author_facet Zuo, Haomin
Li, Yidi
Yang, Luoxiao
Zhang, Xiaofeng
contents While diffusion Multimodal Large Language Models (dMLLMs) have recently achieved remarkable strides in multimodal generation, the development of interpretability mechanisms has lagged behind their architectural evolution. Unlike traditional autoregressive models that produce sequential activations, diffusion-based architectures generate tokens via parallel denoising, resulting in smooth, distributed activation patterns across the entire sequence. Consequently, existing Class Activation Mapping (CAM) methods, which are tailored for local, sequential dependencies, are ill-suited for interpreting these non-autoregressive behaviors. To bridge this gap, we propose Diffusion-CAM, the first interpretability method specifically tailored for dMLLMs. We derive raw activation maps by differentiably probing intermediate representations in the transformer backbone, accordingly capturing both latent features and their class-specific gradients. To address the inherent stochasticity of these raw signals, we incorporate four key modules to resolve spatial ambiguity and mitigate intra-image confounders and redundant token correlations. Extensive experiments demonstrate that Diffusion-CAM significantly outperforms SoTA methods in both localization accuracy and visual fidelity, establishing a new standard for understanding the parallel generation process of diffusion multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11005
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diffusion-CAM: Faithful Visual Explanations for dMLLMs
Zuo, Haomin
Li, Yidi
Yang, Luoxiao
Zhang, Xiaofeng
Artificial Intelligence
While diffusion Multimodal Large Language Models (dMLLMs) have recently achieved remarkable strides in multimodal generation, the development of interpretability mechanisms has lagged behind their architectural evolution. Unlike traditional autoregressive models that produce sequential activations, diffusion-based architectures generate tokens via parallel denoising, resulting in smooth, distributed activation patterns across the entire sequence. Consequently, existing Class Activation Mapping (CAM) methods, which are tailored for local, sequential dependencies, are ill-suited for interpreting these non-autoregressive behaviors. To bridge this gap, we propose Diffusion-CAM, the first interpretability method specifically tailored for dMLLMs. We derive raw activation maps by differentiably probing intermediate representations in the transformer backbone, accordingly capturing both latent features and their class-specific gradients. To address the inherent stochasticity of these raw signals, we incorporate four key modules to resolve spatial ambiguity and mitigate intra-image confounders and redundant token correlations. Extensive experiments demonstrate that Diffusion-CAM significantly outperforms SoTA methods in both localization accuracy and visual fidelity, establishing a new standard for understanding the parallel generation process of diffusion multimodal systems.
title Diffusion-CAM: Faithful Visual Explanations for dMLLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2604.11005