SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Kenan, Li, Yanhong, Qin, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912653286309888
author Tang, Kenan
Li, Yanhong
Qin, Yao
author_facet Tang, Kenan
Li, Yanhong
Qin, Yao
contents Prompt-based models have demonstrated impressive prompt-following capability at image editing tasks. However, the models still struggle with following detailed editing prompts or performing local edits. Specifically, global image quality often deteriorates immediately after a single editing step. To address these challenges, we introduce SPICE, a training-free workflow that accepts arbitrary resolutions and aspect ratios, accurately follows user requirements, and consistently improves image quality during more than 100 editing steps, while keeping the unedited regions intact. By synergizing the strengths of a base diffusion model and a Canny edge ControlNet model, SPICE robustly handles free-form editing instructions from the user. On a challenging realistic image-editing dataset, SPICE quantitatively outperforms state-of-the-art baselines and is consistently preferred by human annotators. We release the workflow implementation for popular diffusion model Web UIs to support further research and artistic exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09697
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow
Tang, Kenan
Li, Yanhong
Qin, Yao
Graphics
Computer Vision and Pattern Recognition
Machine Learning
Prompt-based models have demonstrated impressive prompt-following capability at image editing tasks. However, the models still struggle with following detailed editing prompts or performing local edits. Specifically, global image quality often deteriorates immediately after a single editing step. To address these challenges, we introduce SPICE, a training-free workflow that accepts arbitrary resolutions and aspect ratios, accurately follows user requirements, and consistently improves image quality during more than 100 editing steps, while keeping the unedited regions intact. By synergizing the strengths of a base diffusion model and a Canny edge ControlNet model, SPICE robustly handles free-form editing instructions from the user. On a challenging realistic image-editing dataset, SPICE quantitatively outperforms state-of-the-art baselines and is consistently preferred by human annotators. We release the workflow implementation for popular diffusion model Web UIs to support further research and artistic exploration.
title SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow
topic Graphics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.09697