GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Mingxin, Fan, Ziqian, Wang, Zhaokai, Gu, Leyao, Zhu, Zirun, He, Yiguo, Yang, Yuchen, Tian, Changyao, Zhao, Xiangyu, Liao, Ning, Zhang, Shaofeng, Ren, Qibing, Zhong, Zhihang, Zhou, Xuanhe, Yan, Junchi, Yang, Xue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908882176049152
author Liu, Mingxin
Fan, Ziqian
Wang, Zhaokai
Gu, Leyao
Zhu, Zirun
He, Yiguo
Yang, Yuchen
Tian, Changyao
Zhao, Xiangyu
Liao, Ning
Zhang, Shaofeng
Ren, Qibing
Zhong, Zhihang
Zhou, Xuanhe
Yan, Junchi
Yang, Xue
author_facet Liu, Mingxin
Fan, Ziqian
Wang, Zhaokai
Gu, Leyao
Zhu, Zirun
He, Yiguo
Yang, Yuchen
Tian, Changyao
Zhao, Xiangyu
Liao, Ning
Zhang, Shaofeng
Ren, Qibing
Zhong, Zhihang
Zhou, Xuanhe
Yan, Junchi
Yang, Xue
contents Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain-specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline-informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi-dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in current models under implicit, knowledge-intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline-informed image editing and reasoning. Our benchmark and evaluation code are publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12264
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
Liu, Mingxin
Fan, Ziqian
Wang, Zhaokai
Gu, Leyao
Zhu, Zirun
He, Yiguo
Yang, Yuchen
Tian, Changyao
Zhao, Xiangyu
Liao, Ning
Zhang, Shaofeng
Ren, Qibing
Zhong, Zhihang
Zhou, Xuanhe
Yan, Junchi
Yang, Xue
Computer Vision and Pattern Recognition
Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability under structured, domain-specific constraints. In this work, we introduce GRADE, the first benchmark to assess discipline-informed knowledge and reasoning in image editing. GRADE comprises 520 carefully curated samples across 10 academic domains, spanning from natural science to social science. To support rigorous evaluation, we propose a multi-dimensional evaluation protocol that jointly assesses Discipline Reasoning, Visual Consistency, and Logical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in current models under implicit, knowledge-intensive editing settings, leading to large performance gaps. Beyond quantitative scores, we conduct rigorous analyses and ablations to expose model shortcomings and identify the constraints within disciplinary editing. Together, GRADE pinpoints key directions for the future development of unified multimodal models, advancing the research on discipline-informed image editing and reasoning. Our benchmark and evaluation code are publicly released.
title GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12264