MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Xuehai, Gu, Xiaoling, Liu, Akide, Yuan, Hangjie, Zhang, YiFan, Ma, Jack |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026)
by: Zhang, Huanyu, et al.
Published: (2026)
CompBench: Benchmarking Complex Instruction-guided Image Editing
by: Jia, Bohan, et al.
Published: (2025)
by: Jia, Bohan, et al.
Published: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
by: Wang, Chenglin, et al.
Published: (2025)
by: Wang, Chenglin, et al.
Published: (2025)
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
by: Tian, Xueyun, et al.
Published: (2025)
by: Tian, Xueyun, et al.
Published: (2025)
Multi-Reward as Condition for Instruction-based Image Editing
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Noise Map Guidance: Inversion with Spatial Context for Real Image Editing
by: Cho, Hansam, et al.
Published: (2024)
by: Cho, Hansam, et al.
Published: (2024)
StyleBooth: Image Style Editing with Multimodal Instruction
by: Han, Zhen, et al.
Published: (2024)
by: Han, Zhen, et al.
Published: (2024)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
by: Lin, Yiqi, et al.
Published: (2026)
by: Lin, Yiqi, et al.
Published: (2026)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
Group Relative Attention Guidance for Image Editing
by: Zhang, Xuanpu, et al.
Published: (2025)
by: Zhang, Xuanpu, et al.
Published: (2025)
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
by: Chen, Zhihong, et al.
Published: (2025)
by: Chen, Zhihong, et al.
Published: (2025)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
InstructBrush: Learning Attention-based Instruction Optimization for Image Editing
by: Zhao, Ruoyu, et al.
Published: (2024)
by: Zhao, Ruoyu, et al.
Published: (2024)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
by: Yu, Haoran, et al.
Published: (2025)
by: Yu, Haoran, et al.
Published: (2025)
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
by: Qu, Tianyuan, et al.
Published: (2025)
by: Qu, Tianyuan, et al.
Published: (2025)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
by: Wu, Song, et al.
Published: (2026)
by: Wu, Song, et al.
Published: (2026)
Guidance Free Image Editing via Explicit Conditioning
by: Noroozi, Mehdi, et al.
Published: (2025)
by: Noroozi, Mehdi, et al.
Published: (2025)
Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions
by: Ma, Chenrui, et al.
Published: (2025)
by: Ma, Chenrui, et al.
Published: (2025)
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Debiasing Multimodal Large Language Models via Penalization of Language Priors
by: Zhang, YiFan, et al.
Published: (2024)
by: Zhang, YiFan, et al.
Published: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
by: Jiang, Hong, et al.
Published: (2026)
by: Jiang, Hong, et al.
Published: (2026)
ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
by: Chen, Sherry X., et al.
Published: (2025)
by: Chen, Sherry X., et al.
Published: (2025)
DreamOmni2: Multimodal Instruction-based Editing and Generation
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
Neural-Driven Image Editing
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
by: Yan, Zexuan, et al.
Published: (2025)
by: Yan, Zexuan, et al.
Published: (2025)
Adversarial-Guided Diffusion for Multimodal LLM Attacks
by: Xia, Chengwei, et al.
Published: (2025)
by: Xia, Chengwei, et al.
Published: (2025)
Self-Evolving 3D Scene Generation from a Single Image
by: Zheng, Kaizhi, et al.
Published: (2025)
by: Zheng, Kaizhi, et al.
Published: (2025)
MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
by: Li, Mingsong, et al.
Published: (2025)
by: Li, Mingsong, et al.
Published: (2025)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025)
by: Yang, Siwei, et al.
Published: (2025)
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
by: Cai, Lingling, et al.
Published: (2025)
by: Cai, Lingling, et al.
Published: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
TIE: Revolutionizing Text-based Image Editing for Complex-Prompt Following and High-Fidelity Editing
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
DragScene: Interactive 3D Scene Editing with Single-view Drag Instructions
by: Gu, Chenghao, et al.
Published: (2024)
by: Gu, Chenghao, et al.
Published: (2024)
Similar Items
-
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026) -
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026) -
CompBench: Benchmarking Complex Instruction-guided Image Editing
by: Jia, Bohan, et al.
Published: (2025) -
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025) -
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
by: Wang, Chenglin, et al.
Published: (2025)