RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Tianyu, Ma, Zhiyuan, Wang, Qian, Zhang, Xinyi, Long, Xinwei, Zhou, Bowen
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913100537528320
author Wang, Tianyu
Ma, Zhiyuan
Wang, Qian
Zhang, Xinyi
Long, Xinwei
Zhou, Bowen
author_facet Wang, Tianyu
Ma, Zhiyuan
Wang, Qian
Zhang, Xinyi
Long, Xinwei
Zhou, Bowen
contents Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models still generally struggle with a spatial reasoning dilemma, lacking the ability to accurately capture fine-grained spatial relationships from the prompt and correctly generate scenes with structural integrity. To mitigate this dilemma, we propose RL-RIG, a Reinforcement Learning framework for Reflection-based Image Generation. Our architecture comprises four primary components: Diffuser, Checker, Actor, and Inverse Diffuser, following a Generate-Reflect-Edit paradigm to spark the Chain of Thought reasoning ability in image generation for addressing the dilemma. To equip the model with better intuition over generation trajectories, we further develop Reflection-GRPO to train the VLM Actor for edit prompts and the Image Editor for better image quality under a given prompt, respectively. Unlike traditional approaches that solely produce visually stunning yet structurally unreasonable content, our evaluation metrics prioritize spatial accuracy, utilizing Scene Graph IoU and employing a VLM-as-a-Judge strategy to assess the spatial consistency of generated images on LAION-SG dataset. Experimental results show that RL-RIG outperforms existing state-of-the-art open-source models by up to 11% in terms of controllable and precise spatial reasoning in image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19974
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection
Wang, Tianyu
Ma, Zhiyuan
Wang, Qian
Zhang, Xinyi
Long, Xinwei
Zhou, Bowen
Computer Vision and Pattern Recognition
Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models still generally struggle with a spatial reasoning dilemma, lacking the ability to accurately capture fine-grained spatial relationships from the prompt and correctly generate scenes with structural integrity. To mitigate this dilemma, we propose RL-RIG, a Reinforcement Learning framework for Reflection-based Image Generation. Our architecture comprises four primary components: Diffuser, Checker, Actor, and Inverse Diffuser, following a Generate-Reflect-Edit paradigm to spark the Chain of Thought reasoning ability in image generation for addressing the dilemma. To equip the model with better intuition over generation trajectories, we further develop Reflection-GRPO to train the VLM Actor for edit prompts and the Image Editor for better image quality under a given prompt, respectively. Unlike traditional approaches that solely produce visually stunning yet structurally unreasonable content, our evaluation metrics prioritize spatial accuracy, utilizing Scene Graph IoU and employing a VLM-as-a-Judge strategy to assess the spatial consistency of generated images on LAION-SG dataset. Experimental results show that RL-RIG outperforms existing state-of-the-art open-source models by up to 11% in terms of controllable and precise spatial reasoning in image generation.
title RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19974