Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Xuehai, Shi, Yang, Zhang, Yi-Fan, Zhu, Xuanyu, Wang, Yuran, Dai, Yifan, Liu, Xinyu, Ji, Yiyan, Gu, Xiaoling, Zhang, Yuanxing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917490624299008
author Bai, Xuehai
Shi, Yang
Zhang, Yi-Fan
Zhu, Xuanyu
Wang, Yuran
Dai, Yifan
Liu, Xinyu
Ji, Yiyan
Gu, Xiaoling
Zhang, Yuanxing
author_facet Bai, Xuehai
Shi, Yang
Zhang, Yi-Fan
Zhu, Xuanyu
Wang, Yuran
Dai, Yifan
Liu, Xinyu
Ji, Yiyan
Gu, Xiaoling
Zhang, Yuanxing
contents Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13062
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
Bai, Xuehai
Shi, Yang
Zhang, Yi-Fan
Zhu, Xuanyu
Wang, Yuran
Dai, Yifan
Liu, Xinyu
Ji, Yiyan
Gu, Xiaoling
Zhang, Yuanxing
Computer Vision and Pattern Recognition
Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.
title Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.13062