ZONE: Zero-Shot Instruction-Guided Local Editing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916202126770176 |
|---|---|
| author | Li, Shanglin Zeng, Bohan Feng, Yutang Gao, Sicheng Liu, Xuhui Liu, Jiaming Lin, Li Tang, Xu Hu, Yao Liu, Jianzhuang Zhang, Baochang |
| author_facet | Li, Shanglin Zeng, Bohan Feng, Yutang Gao, Sicheng Liu, Xuhui Liu, Jiaming Lin, Li Tang, Xu Hu, Yao Liu, Jianzhuang Zhang, Baochang |
| contents | Recent advances in vision-language models like Stable Diffusion have shown remarkable power in creative image synthesis and editing.However, most existing text-to-image editing methods encounter two obstacles: First, the text prompt needs to be carefully crafted to achieve good results, which is not intuitive or user-friendly. Second, they are insensitive to local edits and can irreversibly affect non-edited regions, leaving obvious editing traces. To tackle these problems, we propose a Zero-shot instructiON-guided local image Editing approach, termed ZONE. We first convert the editing intent from the user-provided instruction (e.g., "make his tie blue") into specific image editing regions through InstructPix2Pix. We then propose a Region-IoU scheme for precise image layer extraction from an off-the-shelf segment model. We further develop an edge smoother based on FFT for seamless blending between the layer and the image.Our method allows for arbitrary manipulation of a specific region with a single instruction while preserving the rest. Extensive experiments demonstrate that our ZONE achieves remarkable local editing results and user-friendliness, outperforming state-of-the-art methods. Code is available at https://github.com/lsl001006/ZONE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_16794 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ZONE: Zero-Shot Instruction-Guided Local Editing Li, Shanglin Zeng, Bohan Feng, Yutang Gao, Sicheng Liu, Xuhui Liu, Jiaming Lin, Li Tang, Xu Hu, Yao Liu, Jianzhuang Zhang, Baochang Computer Vision and Pattern Recognition Recent advances in vision-language models like Stable Diffusion have shown remarkable power in creative image synthesis and editing.However, most existing text-to-image editing methods encounter two obstacles: First, the text prompt needs to be carefully crafted to achieve good results, which is not intuitive or user-friendly. Second, they are insensitive to local edits and can irreversibly affect non-edited regions, leaving obvious editing traces. To tackle these problems, we propose a Zero-shot instructiON-guided local image Editing approach, termed ZONE. We first convert the editing intent from the user-provided instruction (e.g., "make his tie blue") into specific image editing regions through InstructPix2Pix. We then propose a Region-IoU scheme for precise image layer extraction from an off-the-shelf segment model. We further develop an edge smoother based on FFT for seamless blending between the layer and the image.Our method allows for arbitrary manipulation of a specific region with a single instruction while preserving the rest. Extensive experiments demonstrate that our ZONE achieves remarkable local editing results and user-friendliness, outperforming state-of-the-art methods. Code is available at https://github.com/lsl001006/ZONE. |
| title | ZONE: Zero-Shot Instruction-Guided Local Editing |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2312.16794 |