Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeh, Chun-Hsiao, Wang, Yilin, Zhao, Nanxuan, Zhang, Richard, Li, Yuheng, Ma, Yi, Singh, Krishna Kumar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909677706543104
author Yeh, Chun-Hsiao
Wang, Yilin
Zhao, Nanxuan
Zhang, Richard
Li, Yuheng
Ma, Yi
Singh, Krishna Kumar
author_facet Yeh, Chun-Hsiao
Wang, Yilin
Zhao, Nanxuan
Zhang, Richard
Li, Yuheng
Ma, Yi
Singh, Krishna Kumar
contents Recent diffusion-based image editing methods have significantly advanced text-guided tasks but often struggle to interpret complex, indirect instructions. Moreover, current models frequently suffer from poor identity preservation, unintended edits, or rely heavily on manual masks. To address these challenges, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that effectively bridges user intent with editing model capabilities. X-Planner employs chain-of-thought reasoning to systematically decompose complex instructions into simpler, clear sub-instructions. For each sub-instruction, X-Planner automatically generates precise edit types and segmentation masks, eliminating manual intervention and ensuring localized, identity-preserving edits. Additionally, we propose a novel automated pipeline for generating large-scale data to train X-Planner which achieves state-of-the-art results on both existing benchmarks and our newly introduced complex editing benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05259
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
Yeh, Chun-Hsiao
Wang, Yilin
Zhao, Nanxuan
Zhang, Richard
Li, Yuheng
Ma, Yi
Singh, Krishna Kumar
Computer Vision and Pattern Recognition
Recent diffusion-based image editing methods have significantly advanced text-guided tasks but often struggle to interpret complex, indirect instructions. Moreover, current models frequently suffer from poor identity preservation, unintended edits, or rely heavily on manual masks. To address these challenges, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that effectively bridges user intent with editing model capabilities. X-Planner employs chain-of-thought reasoning to systematically decompose complex instructions into simpler, clear sub-instructions. For each sub-instruction, X-Planner automatically generates precise edit types and segmentation masks, eliminating manual intervention and ensuring localized, identity-preserving edits. Additionally, we propose a novel automated pipeline for generating large-scale data to train X-Planner which achieves state-of-the-art results on both existing benchmarks and our newly introduced complex editing benchmark.
title Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05259