Interleaving Reasoning for Better Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Wenxuan, Chen, Shuang, Xie, Zheyong, Cao, Shaosheng, Tang, Shixiang, Shen, Yufan, Yin, Qingyu, Hu, Wenbo, Wang, Xiaoman, Tang, Yuntian, Qiao, Junbo, Guo, Yue, Hu, Yao, Yin, Zhenfei, Torr, Philip, Cheng, Yu, Ouyang, Wanli, Lin, Shaohui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912577756332032
author Huang, Wenxuan
Chen, Shuang
Xie, Zheyong
Cao, Shaosheng
Tang, Shixiang
Shen, Yufan
Yin, Qingyu
Hu, Wenbo
Wang, Xiaoman
Tang, Yuntian
Qiao, Junbo
Guo, Yue
Hu, Yao
Yin, Zhenfei
Torr, Philip
Cheng, Yu
Ouyang, Wanli
Lin, Shaohui
author_facet Huang, Wenxuan
Chen, Shuang
Xie, Zheyong
Cao, Shaosheng
Tang, Shixiang
Shen, Yufan
Yin, Qingyu
Hu, Wenbo
Wang, Xiaoman
Tang, Yuntian
Qiao, Junbo
Guo, Yue
Hu, Yao
Yin, Zhenfei
Torr, Philip
Cheng, Yu
Ouyang, Wanli
Lin, Shaohui
contents Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly couple comprehension with generation such as GPT-4o. Motivated by recent advances in interleaving reasoning, we explore whether such reasoning can further improve Text-to-Image (T2I) generation. We introduce Interleaving Reasoning Generation (IRG), a framework that alternates between text-based thinking and image synthesis: the model first produces a text-based thinking to guide an initial image, then reflects on the result to refine fine-grained details, visual quality, and aesthetics while preserving semantics. To train IRG effectively, we propose Interleaving Reasoning Generation Learning (IRGL), which targets two sub-goals: (1) strengthening the initial think-and-generate stage to establish core content and base quality, and (2) enabling high-quality textual reflection and faithful implementation of those refinements in a subsequent image. We curate IRGL-300K, a dataset organized into six decomposed learning modes that jointly cover learning text-based thinking, and full thinking-image trajectories. Starting from a unified foundation model that natively emits interleaved text-image outputs, our two-stage training first builds robust thinking and reflection, then efficiently tunes the IRG pipeline in the full thinking-image trajectory data. Extensive experiments show SoTA performance, yielding absolute gains of 5-10 points on GenEval, WISE, TIIF, GenAI-Bench, and OneIG-EN, alongside substantial improvements in visual quality and fine-grained fidelity. The code, model weights and datasets will be released in: https://github.com/Osilly/Interleaving-Reasoning-Generation .
format Preprint
id arxiv_https___arxiv_org_abs_2509_06945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interleaving Reasoning for Better Text-to-Image Generation
Huang, Wenxuan
Chen, Shuang
Xie, Zheyong
Cao, Shaosheng
Tang, Shixiang
Shen, Yufan
Yin, Qingyu
Hu, Wenbo
Wang, Xiaoman
Tang, Yuntian
Qiao, Junbo
Guo, Yue
Hu, Yao
Yin, Zhenfei
Torr, Philip
Cheng, Yu
Ouyang, Wanli
Lin, Shaohui
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly couple comprehension with generation such as GPT-4o. Motivated by recent advances in interleaving reasoning, we explore whether such reasoning can further improve Text-to-Image (T2I) generation. We introduce Interleaving Reasoning Generation (IRG), a framework that alternates between text-based thinking and image synthesis: the model first produces a text-based thinking to guide an initial image, then reflects on the result to refine fine-grained details, visual quality, and aesthetics while preserving semantics. To train IRG effectively, we propose Interleaving Reasoning Generation Learning (IRGL), which targets two sub-goals: (1) strengthening the initial think-and-generate stage to establish core content and base quality, and (2) enabling high-quality textual reflection and faithful implementation of those refinements in a subsequent image. We curate IRGL-300K, a dataset organized into six decomposed learning modes that jointly cover learning text-based thinking, and full thinking-image trajectories. Starting from a unified foundation model that natively emits interleaved text-image outputs, our two-stage training first builds robust thinking and reflection, then efficiently tunes the IRG pipeline in the full thinking-image trajectory data. Extensive experiments show SoTA performance, yielding absolute gains of 5-10 points on GenEval, WISE, TIIF, GenAI-Bench, and OneIG-EN, alongside substantial improvements in visual quality and fine-grained fidelity. The code, model weights and datasets will be released in: https://github.com/Osilly/Interleaving-Reasoning-Generation .
title Interleaving Reasoning for Better Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.06945