Watch Your Steps: Local Image and Scene Editing by Text Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mirzaei, Ashkan, Aumentado-Armstrong, Tristan, Brubaker, Marcus A., Kelly, Jonathan, Levinshtein, Alex, Derpanis, Konstantinos G., Gilitschenski, Igor
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910510356627456
author Mirzaei, Ashkan
Aumentado-Armstrong, Tristan
Brubaker, Marcus A.
Kelly, Jonathan
Levinshtein, Alex
Derpanis, Konstantinos G.
Gilitschenski, Igor
author_facet Mirzaei, Ashkan
Aumentado-Armstrong, Tristan
Brubaker, Marcus A.
Kelly, Jonathan
Levinshtein, Alex
Derpanis, Konstantinos G.
Gilitschenski, Igor
contents Denoising diffusion models have enabled high-quality image generation and editing. We present a method to localize the desired edit region implicit in a text instruction. We leverage InstructPix2Pix (IP2P) and identify the discrepancy between IP2P predictions with and without the instruction. This discrepancy is referred to as the relevance map. The relevance map conveys the importance of changing each pixel to achieve the edits, and is used to to guide the modifications. This guidance ensures that the irrelevant pixels remain unchanged. Relevance maps are further used to enhance the quality of text-guided editing of 3D scenes in the form of neural radiance fields. A field is trained on relevance maps of training views, denoted as the relevance field, defining the 3D region within which modifications should be made. We perform iterative updates on the training views guided by rendered relevance maps from the relevance field. Our method achieves state-of-the-art performance on both image and NeRF editing tasks. Project page: https://ashmrz.github.io/WatchYourSteps/
format Preprint
id arxiv_https___arxiv_org_abs_2308_08947
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Watch Your Steps: Local Image and Scene Editing by Text Instructions
Mirzaei, Ashkan
Aumentado-Armstrong, Tristan
Brubaker, Marcus A.
Kelly, Jonathan
Levinshtein, Alex
Derpanis, Konstantinos G.
Gilitschenski, Igor
Computer Vision and Pattern Recognition
Denoising diffusion models have enabled high-quality image generation and editing. We present a method to localize the desired edit region implicit in a text instruction. We leverage InstructPix2Pix (IP2P) and identify the discrepancy between IP2P predictions with and without the instruction. This discrepancy is referred to as the relevance map. The relevance map conveys the importance of changing each pixel to achieve the edits, and is used to to guide the modifications. This guidance ensures that the irrelevant pixels remain unchanged. Relevance maps are further used to enhance the quality of text-guided editing of 3D scenes in the form of neural radiance fields. A field is trained on relevance maps of training views, denoted as the relevance field, defining the 3D region within which modifications should be made. We perform iterative updates on the training views guided by rendered relevance maps from the relevance field. Our method achieves state-of-the-art performance on both image and NeRF editing tasks. Project page: https://ashmrz.github.io/WatchYourSteps/
title Watch Your Steps: Local Image and Scene Editing by Text Instructions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.08947