R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Kaijie, Lin, Zihao, Xu, Zhiyang, Shen, Ying, Yao, Yuguang, Rimchala, Joy, Zhang, Jiaxin, Huang, Lifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918038188589056
author Chen, Kaijie
Lin, Zihao
Xu, Zhiyang
Shen, Ying
Yao, Yuguang
Rimchala, Joy
Zhang, Jiaxin
Huang, Lifu
author_facet Chen, Kaijie
Lin, Zihao
Xu, Zhiyang
Shen, Ying
Yao, Yuguang
Rimchala, Joy
Zhang, Jiaxin
Huang, Lifu
contents Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and commonsense concepts. While recent T2I models have made impressive progress in producing photorealistic images, their reasoning capability remains underdeveloped and insufficiently evaluated. To bridge this gap, we introduce R2I-Bench, a comprehensive benchmark specifically designed to rigorously assess reasoning-driven T2I generation. R2I-Bench comprises meticulously curated data instances, spanning core reasoning categories, including commonsense, mathematical, logical, compositional, numerical, causal, and concept mixing. To facilitate fine-grained evaluation, we design R2IScore, a QA-style metric based on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality. Extensive experiments with 16 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems. Project Page: https://r2i-bench.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2505_23493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
Chen, Kaijie
Lin, Zihao
Xu, Zhiyang
Shen, Ying
Yao, Yuguang
Rimchala, Joy
Zhang, Jiaxin
Huang, Lifu
Computer Vision and Pattern Recognition
Computation and Language
Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and commonsense concepts. While recent T2I models have made impressive progress in producing photorealistic images, their reasoning capability remains underdeveloped and insufficiently evaluated. To bridge this gap, we introduce R2I-Bench, a comprehensive benchmark specifically designed to rigorously assess reasoning-driven T2I generation. R2I-Bench comprises meticulously curated data instances, spanning core reasoning categories, including commonsense, mathematical, logical, compositional, numerical, causal, and concept mixing. To facilitate fine-grained evaluation, we design R2IScore, a QA-style metric based on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality. Extensive experiments with 16 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems. Project Page: https://r2i-bench.github.io
title R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.23493