GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fang, Rongyao, Duan, Chengqi, Wang, Kun, Huang, Linjiang, Li, Hao, Yan, Shilin, Tian, Hao, Zeng, Xingyu, Zhao, Rui, Dai, Jifeng, Liu, Xihui, Li, Hongsheng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912273823432704
author Fang, Rongyao
Duan, Chengqi
Wang, Kun
Huang, Linjiang
Li, Hao
Yan, Shilin
Tian, Hao
Zeng, Xingyu
Zhao, Rui
Dai, Jifeng
Liu, Xihui
Li, Hongsheng
author_facet Fang, Rongyao
Duan, Chengqi
Wang, Kun
Huang, Linjiang
Li, Hao
Yan, Shilin
Tian, Hao
Zeng, Xingyu
Zhao, Rui
Dai, Jifeng
Liu, Xihui
Li, Hongsheng
contents Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Generation Chain-of-Thought (GoT), a novel paradigm that enables generation and editing through an explicit language reasoning process before outputting images. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over 9M samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. To facilitate future research, we make our datasets, code, and pretrained models publicly available at https://github.com/rongyaofang/GoT.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
Fang, Rongyao
Duan, Chengqi
Wang, Kun
Huang, Linjiang
Li, Hao
Yan, Shilin
Tian, Hao
Zeng, Xingyu
Zhao, Rui
Dai, Jifeng
Liu, Xihui
Li, Hongsheng
Computer Vision and Pattern Recognition
Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Generation Chain-of-Thought (GoT), a novel paradigm that enables generation and editing through an explicit language reasoning process before outputting images. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over 9M samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. To facilitate future research, we make our datasets, code, and pretrained models publicly available at https://github.com/rongyaofang/GoT.
title GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10639