CompBench: Benchmarking Complex Instruction-guided Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Bohan, Huang, Wenxuan, Tang, Yuntian, Qiao, Junbo, Liao, Jincheng, Cao, Shaosheng, Zhao, Fei, Feng, Zhaopeng, Gu, Zhouhong, Yin, Zhenfei, Bai, Lei, Ouyang, Wanli, Chen, Lin, Hu, Yao, Wang, Zihan, Xie, Yuan, Lin, Shaohui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
by: Sun, Kaiyue, et al.
Published: (2024)
by: Sun, Kaiyue, et al.
Published: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023)
by: Huang, Kaiyi, et al.
Published: (2023)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
by: Zhan, Xiaoyu, et al.
Published: (2025)
by: Zhan, Xiaoyu, et al.
Published: (2025)
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2026)
by: Huang, Wenxuan, et al.
Published: (2026)
Dynamic Contrastive Knowledge Distillation for Efficient Image Restoration
by: Zhou, Yunshuai, et al.
Published: (2024)
by: Zhou, Yunshuai, et al.
Published: (2024)
Allo{SR}$^2$: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
by: Dong, Qin, et al.
Published: (2025)
by: Dong, Qin, et al.
Published: (2025)
From Static Spectra to Operando Infrared Dynamics: Physics Informed Flow Modeling and a Benchmark
by: Ye, Shuquan, et al.
Published: (2026)
by: Ye, Shuquan, et al.
Published: (2026)
Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models
by: Liu, Dingning, et al.
Published: (2024)
by: Liu, Dingning, et al.
Published: (2024)
ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation
by: Zhou, Shizhe, et al.
Published: (2026)
by: Zhou, Shizhe, et al.
Published: (2026)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
by: Zeng, Yu, et al.
Published: (2026)
by: Zeng, Yu, et al.
Published: (2026)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
Intelligent event‐triggered filtering for interval type‐2 fuzzy singular systems
by: Xuan Zhao, et al.
Published: (2025)
by: Xuan Zhao, et al.
Published: (2025)
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression
by: Tang, Yuntian, et al.
Published: (2026)
by: Tang, Yuntian, et al.
Published: (2026)
MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
by: Li, Mingsong, et al.
Published: (2025)
by: Li, Mingsong, et al.
Published: (2025)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
by: Li, Bangyan, et al.
Published: (2025)
by: Li, Bangyan, et al.
Published: (2025)
MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
by: Chen, Yifan, et al.
Published: (2026)
by: Chen, Yifan, et al.
Published: (2026)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization
by: Xu, Kaiyi, et al.
Published: (2025)
by: Xu, Kaiyi, et al.
Published: (2025)
ReconMOST: Multi-Layer Sea Temperature Reconstruction with Observations-Guided Diffusion
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
by: Jiang, Pengfei, et al.
Published: (2024)
by: Jiang, Pengfei, et al.
Published: (2024)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
by: Shen, Hao, et al.
Published: (2025)
by: Shen, Hao, et al.
Published: (2025)
Smartphone User Profiling Based on Multichannel Features Extracted From Installed App List
by: Fei Wang, et al.
Published: (2025)
by: Fei Wang, et al.
Published: (2025)
Quantitative convergence guarantees for the mean-field dispersion process
by: Cao, Fei, et al.
Published: (2024)
by: Cao, Fei, et al.
Published: (2024)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
by: Xiao, Ruixuan, et al.
Published: (2024)
by: Xiao, Ruixuan, et al.
Published: (2024)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
ZONE: Zero-Shot Instruction-Guided Local Editing
by: Li, Shanglin, et al.
Published: (2023)
by: Li, Shanglin, et al.
Published: (2023)
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
by: Xue, Xiangyuan, et al.
Published: (2026)
by: Xue, Xiangyuan, et al.
Published: (2026)
Flow-OPD: On-Policy Distillation for Flow Matching Models
by: Fang, Zhen, et al.
Published: (2026)
by: Fang, Zhen, et al.
Published: (2026)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
InstructionBench: An Instructional Video Understanding Benchmark
by: Wei, Haiwan, et al.
Published: (2025)
by: Wei, Haiwan, et al.
Published: (2025)
Efficiently Quantifying and Mitigating Ripple Effects in Model Editing
by: Wang, Jianchen, et al.
Published: (2024)
by: Wang, Jianchen, et al.
Published: (2024)
UniComp: Rethinking Video Compression Through Informational Uniqueness
by: Yuan, Chao, et al.
Published: (2025)
by: Yuan, Chao, et al.
Published: (2025)
Understanding the Collapse of LLMs in Model Editing
by: Yang, Wanli, et al.
Published: (2024)
by: Yang, Wanli, et al.
Published: (2024)
Regularity and long-time behavior of global weak solutions to a coupled Cahn-Hilliard system: the off-critical case
by: Ouyang, Bohan
Published: (2024)
by: Ouyang, Bohan
Published: (2024)
Similar Items
-
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
by: Sun, Kaiyue, et al.
Published: (2024) -
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
by: Kil, Jihyung, et al.
Published: (2024) -
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023) -
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
by: Zhan, Xiaoyu, et al.
Published: (2025) -
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025)