Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.16707 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913853115203584 |
|---|---|
| author | Wu, Yongliang Li, Zonghui Hu, Xinting Ye, Xinyu Zeng, Xianfang Yu, Gang Zhu, Wenbo Schiele, Bernt Yang, Ming-Hsuan Yang, Xu |
| author_facet | Wu, Yongliang Li, Zonghui Hu, Xinting Ye, Xinyu Zeng, Xianfang Yu, Gang Zhu, Wenbo Schiele, Bernt Yang, Ming-Hsuan Yang, Xu |
| contents | Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16707 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models Wu, Yongliang Li, Zonghui Hu, Xinting Ye, Xinyu Zeng, Xianfang Yu, Gang Zhu, Wenbo Schiele, Bernt Yang, Ming-Hsuan Yang, Xu Computer Vision and Pattern Recognition Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, we introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on 10 state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems. |
| title | KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.16707 |