Unified Thinker: A General Reasoning Modular Core for Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Sashuai, Zhou, Qiang, Hu, Jijin, Yang, Hanqing, Cao, Yue, Ma, Junpeng, Ma, Yinchao, Song, Jun, Ge, Tiezheng, Yu, Cheng, Zheng, Bo, Zhao, Zhou
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911564101058560
author Zhou, Sashuai
Zhou, Qiang
Hu, Jijin
Yang, Hanqing
Cao, Yue
Ma, Junpeng
Ma, Yinchao
Song, Jun
Ge, Tiezheng
Yu, Cheng
Zheng, Bo
Zhao, Zhou
author_facet Zhou, Sashuai
Zhou, Qiang
Hu, Jijin
Yang, Hanqing
Cao, Yue
Ma, Junpeng
Ma, Yinchao
Song, Jun
Ge, Tiezheng
Yu, Cheng
Zheng, Bo
Zhao, Zhou
contents Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--execution gap. Meanwhile, closed-source systems (e.g., Nano Banana) have demonstrated strong reasoning-driven image generation, highlighting a substantial gap to current open-source models. We argue that closing this gap requires not merely better visual generators, but executable reasoning: decomposing high-level intents into grounded, verifiable plans that directly steer the generative process. To this end, we propose Unified Thinker, a task-agnostic reasoning architecture for general image generation, designed as a unified planning core that can plug into diverse generators and workflows. Unified Thinker decouples a dedicated Thinker from the image Generator, enabling modular upgrades of reasoning without retraining the entire generative model. We further introduce a two-stage training paradigm: we first build a structured planning interface for the Thinker, then apply reinforcement learning to ground its policy in pixel-level feedback, encouraging plans that optimize visual correctness over textual plausibility. Extensive experiments on text-to-image generation and image editing show that Unified Thinker substantially improves image reasoning and generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03127
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unified Thinker: A General Reasoning Modular Core for Image Generation
Zhou, Sashuai
Zhou, Qiang
Hu, Jijin
Yang, Hanqing
Cao, Yue
Ma, Junpeng
Ma, Yinchao
Song, Jun
Ge, Tiezheng
Yu, Cheng
Zheng, Bo
Zhao, Zhou
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--execution gap. Meanwhile, closed-source systems (e.g., Nano Banana) have demonstrated strong reasoning-driven image generation, highlighting a substantial gap to current open-source models. We argue that closing this gap requires not merely better visual generators, but executable reasoning: decomposing high-level intents into grounded, verifiable plans that directly steer the generative process. To this end, we propose Unified Thinker, a task-agnostic reasoning architecture for general image generation, designed as a unified planning core that can plug into diverse generators and workflows. Unified Thinker decouples a dedicated Thinker from the image Generator, enabling modular upgrades of reasoning without retraining the entire generative model. We further introduce a two-stage training paradigm: we first build a structured planning interface for the Thinker, then apply reinforcement learning to ground its policy in pixel-level feedback, encouraging plans that optimize visual correctness over textual plausibility. Extensive experiments on text-to-image generation and image editing show that Unified Thinker substantially improves image reasoning and generation quality.
title Unified Thinker: A General Reasoning Modular Core for Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.03127