CompBench: Benchmarking Complex Instruction-guided Image Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jia, Bohan, Huang, Wenxuan, Tang, Yuntian, Qiao, Junbo, Liao, Jincheng, Cao, Shaosheng, Zhao, Fei, Feng, Zhaopeng, Gu, Zhouhong, Yin, Zhenfei, Bai, Lei, Ouyang, Wanli, Chen, Lin, Hu, Yao, Wang, Zihan, Xie, Yuan, Lin, Shaohui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914423437787136
author Jia, Bohan
Huang, Wenxuan
Tang, Yuntian
Qiao, Junbo
Liao, Jincheng
Cao, Shaosheng
Zhao, Fei
Feng, Zhaopeng
Gu, Zhouhong
Yin, Zhenfei
Bai, Lei
Ouyang, Wanli
Chen, Lin
Zhao, Fei
Hu, Yao
Wang, Zihan
Xie, Yuan
Lin, Shaohui
author_facet Jia, Bohan
Huang, Wenxuan
Tang, Yuntian
Qiao, Junbo
Liao, Jincheng
Cao, Shaosheng
Zhao, Fei
Feng, Zhaopeng
Gu, Zhouhong
Yin, Zhenfei
Bai, Lei
Ouyang, Wanli
Chen, Lin
Zhao, Fei
Hu, Yao
Wang, Zihan
Xie, Yuan
Lin, Shaohui
contents While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap, we introduce CompBench, a large-scale benchmark specifically designed for complex instruction-guided image editing. CompBench features challenging editing scenarios that incorporate fine-grained instruction following, spatial and contextual reasoning, thereby enabling comprehensive evaluation of image editing models' precise manipulation capabilities. To construct CompBench, we propose an MLLM-human collaborative framework with tailored task pipelines. Furthermore, we propose an instruction decoupling strategy that disentangles editing intents into four key dimensions: location, appearance, dynamics, and objects, ensuring closer alignment between instructions and complex editing requirements. Extensive evaluations reveal that CompBench exposes fundamental limitations of current image editing models and provides critical insights for the development of next-generation instruction-guided image editing systems. Our project page is available at https://comp-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12200
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CompBench: Benchmarking Complex Instruction-guided Image Editing
Jia, Bohan
Huang, Wenxuan
Tang, Yuntian
Qiao, Junbo
Liao, Jincheng
Cao, Shaosheng
Zhao, Fei
Feng, Zhaopeng
Gu, Zhouhong
Yin, Zhenfei
Bai, Lei
Ouyang, Wanli
Chen, Lin
Zhao, Fei
Hu, Yao
Wang, Zihan
Xie, Yuan
Lin, Shaohui
Computer Vision and Pattern Recognition
While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap, we introduce CompBench, a large-scale benchmark specifically designed for complex instruction-guided image editing. CompBench features challenging editing scenarios that incorporate fine-grained instruction following, spatial and contextual reasoning, thereby enabling comprehensive evaluation of image editing models' precise manipulation capabilities. To construct CompBench, we propose an MLLM-human collaborative framework with tailored task pipelines. Furthermore, we propose an instruction decoupling strategy that disentangles editing intents into four key dimensions: location, appearance, dynamics, and objects, ensuring closer alignment between instructions and complex editing requirements. Extensive evaluations reveal that CompBench exposes fundamental limitations of current image editing models and provides critical insights for the development of next-generation instruction-guided image editing systems. Our project page is available at https://comp-bench.github.io/.
title CompBench: Benchmarking Complex Instruction-guided Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12200