Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Yuanhuiyi, Wong, Chi Kit, Liao, Chenfei, Jiang, Lutao, Zheng, Xu, Lu, Zexin, Zhang, Linfeng, Hu, Xuming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914055657095168
author Lyu, Yuanhuiyi
Wong, Chi Kit
Liao, Chenfei
Jiang, Lutao
Zheng, Xu
Lu, Zexin
Zhang, Linfeng
Hu, Xuming
author_facet Lyu, Yuanhuiyi
Wong, Chi Kit
Liao, Chenfei
Jiang, Lutao
Zheng, Xu
Lu, Zexin
Zhang, Linfeng
Hu, Xuming
contents Recent works have made notable advancements in enhancing unified models for text-to-image generation through the Chain-of-Thought (CoT). However, these reasoning methods separate the processes of understanding and generation, which limits their ability to guide the reasoning of unified models in addressing the deficiencies of their generative capabilities. To this end, we propose a novel reasoning framework for unified models, Understanding-in-Generation (UiG), which harnesses the robust understanding capabilities of unified models to reinforce their performance in image generation. The core insight of our UiG is to integrate generative guidance by the strong understanding capabilities during the reasoning process, thereby mitigating the limitations of generative abilities. To achieve this, we introduce "Image Editing" as a bridge to infuse understanding into the generation process. Initially, we verify the generated image and incorporate the understanding of unified models into the editing instructions. Subsequently, we enhance the generated image step by step, gradually infusing the understanding into the generation process. Our UiG framework demonstrates a significant performance improvement in text-to-image generation over existing text-to-image reasoning methods, e.g., a 3.92% gain on the long prompt setting of the TIIF benchmark. The project code: https://github.com/QC-LY/UiG
format Preprint
id arxiv_https___arxiv_org_abs_2509_18639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
Lyu, Yuanhuiyi
Wong, Chi Kit
Liao, Chenfei
Jiang, Lutao
Zheng, Xu
Lu, Zexin
Zhang, Linfeng
Hu, Xuming
Computer Vision and Pattern Recognition
Recent works have made notable advancements in enhancing unified models for text-to-image generation through the Chain-of-Thought (CoT). However, these reasoning methods separate the processes of understanding and generation, which limits their ability to guide the reasoning of unified models in addressing the deficiencies of their generative capabilities. To this end, we propose a novel reasoning framework for unified models, Understanding-in-Generation (UiG), which harnesses the robust understanding capabilities of unified models to reinforce their performance in image generation. The core insight of our UiG is to integrate generative guidance by the strong understanding capabilities during the reasoning process, thereby mitigating the limitations of generative abilities. To achieve this, we introduce "Image Editing" as a bridge to infuse understanding into the generation process. Initially, we verify the generated image and incorporate the understanding of unified models into the editing instructions. Subsequently, we enhance the generated image step by step, gradually infusing the understanding into the generation process. Our UiG framework demonstrates a significant performance improvement in text-to-image generation over existing text-to-image reasoning methods, e.g., a 3.92% gain on the long prompt setting of the TIIF benchmark. The project code: https://github.com/QC-LY/UiG
title Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18639