DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Zefeng, Qu, Xiaoye, Li, Yafu, Zhu, Tong, Huang, Siyuan, Cheng, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918266167885824
author He, Zefeng
Qu, Xiaoye
Li, Yafu
Zhu, Tong
Huang, Siyuan
Cheng, Yu
author_facet He, Zefeng
Qu, Xiaoye
Li, Yafu
Zhu, Tong
Huang, Siyuan
Cheng, Yu
contents While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leading to suboptimal performance in complex long-horizon, vision-centric tasks. In this paper, we establish a novel Generative Multimodal Reasoning paradigm and introduce DiffThinker, a diffusion-based reasoning framework. Conceptually, DiffThinker reformulates multimodal reasoning as a native generative image-to-image task, achieving superior logical consistency and spatial precision in vision-centric tasks. We perform a systematic comparison between DiffThinker and MLLMs, providing the first in-depth investigation into the intrinsic characteristics of this paradigm, revealing four core properties: efficiency, controllability, native parallelism, and collaboration. Extensive experiments across four domains (sequential planning, combinatorial optimization, constraint satisfaction, and spatial configuration) demonstrate that DiffThinker significantly outperforms leading closed source models including GPT-5 (+314.2\%) and Gemini-3-Flash (+111.6\%), as well as the fine-tuned Qwen3-VL-32B baseline (+39.0\%), highlighting generative multimodal reasoning as a promising approach for vision-centric reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24165
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
He, Zefeng
Qu, Xiaoye
Li, Yafu
Zhu, Tong
Huang, Siyuan
Cheng, Yu
Computer Vision and Pattern Recognition
While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leading to suboptimal performance in complex long-horizon, vision-centric tasks. In this paper, we establish a novel Generative Multimodal Reasoning paradigm and introduce DiffThinker, a diffusion-based reasoning framework. Conceptually, DiffThinker reformulates multimodal reasoning as a native generative image-to-image task, achieving superior logical consistency and spatial precision in vision-centric tasks. We perform a systematic comparison between DiffThinker and MLLMs, providing the first in-depth investigation into the intrinsic characteristics of this paradigm, revealing four core properties: efficiency, controllability, native parallelism, and collaboration. Extensive experiments across four domains (sequential planning, combinatorial optimization, constraint satisfaction, and spatial configuration) demonstrate that DiffThinker significantly outperforms leading closed source models including GPT-5 (+314.2\%) and Gemini-3-Flash (+111.6\%), as well as the fine-tuned Qwen3-VL-32B baseline (+39.0\%), highlighting generative multimodal reasoning as a promising approach for vision-centric reasoning.
title DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.24165