GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hongxiang, Li, Yaowei, Lin, Bin, Niu, Yuwei, Yang, Yuhang, Huang, Xiaoshuang, Cai, Jiayin, Jiang, Xiaolong, Hu, Yao, Chen, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911533292847104
author Li, Hongxiang
Li, Yaowei
Lin, Bin
Niu, Yuwei
Yang, Yuhang
Huang, Xiaoshuang
Cai, Jiayin
Jiang, Xiaolong
Hu, Yao
Chen, Long
author_facet Li, Hongxiang
Li, Yaowei
Lin, Bin
Niu, Yuwei
Yang, Yuhang
Huang, Xiaoshuang
Cai, Jiayin
Jiang, Xiaolong
Hu, Yao
Chen, Long
contents Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the alignment between understanding and generation, and their generalization potential in complex visual tasks. To this end, we introduce GIR-Bench, a comprehensive benchmark that evaluates unified models across three complementary perspectives. Firstly, we investigate understanding-generation consistency (GIR-Bench-UGC), asking whether models can consistently leverage the same knowledge in both understanding and generation tasks. Secondly, we investigate whether models can perform reasoning-centric text-to-image generation that requires applying logical constraints and implicit knowledge to generate faithful visual content (GIR-Bench-T2I). Thirdly, we evaluate whether models can handle multi-step reasoning in editing (GIR-Bench-Edit). For each subset, we carefully design different task-specific evaluation pipelines tailored for each task. This enables fine-grained and interpretable evaluation while mitigating biases from the prevalent MLLM-as-a-Judge paradigm. Extensive ablations over various unified models and generation-only systems have shown that: Although unified models are more capable of reasoning-driven visual tasks, they still exhibit a persistent gap between understanding and generation. The data and code for GIR-Bench are available at https://github.com/HKUST-LongGroup/GIR-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11026
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
Li, Hongxiang
Li, Yaowei
Lin, Bin
Niu, Yuwei
Yang, Yuhang
Huang, Xiaoshuang
Cai, Jiayin
Jiang, Xiaolong
Hu, Yao
Chen, Long
Computer Vision and Pattern Recognition
Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the alignment between understanding and generation, and their generalization potential in complex visual tasks. To this end, we introduce GIR-Bench, a comprehensive benchmark that evaluates unified models across three complementary perspectives. Firstly, we investigate understanding-generation consistency (GIR-Bench-UGC), asking whether models can consistently leverage the same knowledge in both understanding and generation tasks. Secondly, we investigate whether models can perform reasoning-centric text-to-image generation that requires applying logical constraints and implicit knowledge to generate faithful visual content (GIR-Bench-T2I). Thirdly, we evaluate whether models can handle multi-step reasoning in editing (GIR-Bench-Edit). For each subset, we carefully design different task-specific evaluation pipelines tailored for each task. This enables fine-grained and interpretable evaluation while mitigating biases from the prevalent MLLM-as-a-Judge paradigm. Extensive ablations over various unified models and generation-only systems have shown that: Although unified models are more capable of reasoning-driven visual tasks, they still exhibit a persistent gap between understanding and generation. The data and code for GIR-Bench are available at https://github.com/HKUST-LongGroup/GIR-Bench.
title GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11026