ThinkGen: Generalized Thinking for Visual Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiao, Siyu, Lin, Yiheng, Zhong, Yujie, She, Qi, Zhou, Wei, Lan, Xiaohan, Huang, Zilong, Yu, Fei, Yu, Yingchen, Zhao, Yunqing, Zhao, Yao, Wei, Yunchao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915698535563264
author Jiao, Siyu
Lin, Yiheng
Zhong, Yujie
She, Qi
Zhou, Wei
Lan, Xiaohan
Huang, Zilong
Yu, Fei
Yu, Yingchen
Zhao, Yunqing
Zhao, Yao
Wei, Yunchao
author_facet Jiao, Siyu
Lin, Yiheng
Zhong, Yujie
She, Qi
Zhou, Wei
Lan, Xiaohan
Huang, Zilong
Yu, Fei
Yu, Yingchen
Zhao, Yunqing
Zhao, Yao
Wei, Yunchao
contents Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However, its extension to generation tasks remains nascent and limited by scenario-specific mechanisms that hinder generalization and adaptation. In this work, we present ThinkGen, the first think-driven visual generation framework that explicitly leverages MLLM's CoT reasoning in various generation scenarios. ThinkGen employs a decoupled architecture comprising a pretrained MLLM and a Diffusion Transformer (DiT), wherein the MLLM generates tailored instructions based on user intent, and DiT produces high-quality images guided by these instructions. We further propose a separable GRPO-based training paradigm (SepGRPO), alternating reinforcement learning between the MLLM and DiT modules. This flexible design enables joint training across diverse datasets, facilitating effective CoT reasoning for a wide range of generative scenarios. Extensive experiments demonstrate that ThinkGen achieves robust, state-of-the-art performance across multiple generation benchmarks. Code is available: https://github.com/jiaosiyuu/ThinkGen
format Preprint
id arxiv_https___arxiv_org_abs_2512_23568
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ThinkGen: Generalized Thinking for Visual Generation
Jiao, Siyu
Lin, Yiheng
Zhong, Yujie
She, Qi
Zhou, Wei
Lan, Xiaohan
Huang, Zilong
Yu, Fei
Yu, Yingchen
Zhao, Yunqing
Zhao, Yao
Wei, Yunchao
Computer Vision and Pattern Recognition
Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However, its extension to generation tasks remains nascent and limited by scenario-specific mechanisms that hinder generalization and adaptation. In this work, we present ThinkGen, the first think-driven visual generation framework that explicitly leverages MLLM's CoT reasoning in various generation scenarios. ThinkGen employs a decoupled architecture comprising a pretrained MLLM and a Diffusion Transformer (DiT), wherein the MLLM generates tailored instructions based on user intent, and DiT produces high-quality images guided by these instructions. We further propose a separable GRPO-based training paradigm (SepGRPO), alternating reinforcement learning between the MLLM and DiT modules. This flexible design enables joint training across diverse datasets, facilitating effective CoT reasoning for a wide range of generative scenarios. Extensive experiments demonstrate that ThinkGen achieves robust, state-of-the-art performance across multiple generation benchmarks. Code is available: https://github.com/jiaosiyuu/ThinkGen
title ThinkGen: Generalized Thinking for Visual Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23568