Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Xingang, Tewari, Ayush, Leimkühler, Thomas, Liu, Lingjie, Meka, Abhimitra, Theobalt, Christian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909258284531712
author Pan, Xingang
Tewari, Ayush
Leimkühler, Thomas
Liu, Lingjie
Meka, Abhimitra
Theobalt, Christian
author_facet Pan, Xingang
Tewari, Ayush
Leimkühler, Thomas
Liu, Lingjie
Meka, Abhimitra
Theobalt, Christian
contents Synthesizing visual content that meets users' needs often requires flexible and precise controllability of the pose, shape, expression, and layout of the generated objects. Existing approaches gain controllability of generative adversarial networks (GANs) via manually annotated training data or a prior 3D model, which often lack flexibility, precision, and generality. In this work, we study a powerful yet much less explored way of controlling GANs, that is, to "drag" any points of the image to precisely reach target points in a user-interactive manner, as shown in Fig.1. To achieve this, we propose DragGAN, which consists of two main components: 1) a feature-based motion supervision that drives the handle point to move towards the target position, and 2) a new point tracking approach that leverages the discriminative generator features to keep localizing the position of the handle points. Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc. As these manipulations are performed on the learned generative image manifold of a GAN, they tend to produce realistic outputs even for challenging scenarios such as hallucinating occluded content and deforming shapes that consistently follow the object's rigidity. Both qualitative and quantitative comparisons demonstrate the advantage of DragGAN over prior approaches in the tasks of image manipulation and point tracking. We also showcase the manipulation of real images through GAN inversion.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10973
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
Pan, Xingang
Tewari, Ayush
Leimkühler, Thomas
Liu, Lingjie
Meka, Abhimitra
Theobalt, Christian
Computer Vision and Pattern Recognition
Graphics
Synthesizing visual content that meets users' needs often requires flexible and precise controllability of the pose, shape, expression, and layout of the generated objects. Existing approaches gain controllability of generative adversarial networks (GANs) via manually annotated training data or a prior 3D model, which often lack flexibility, precision, and generality. In this work, we study a powerful yet much less explored way of controlling GANs, that is, to "drag" any points of the image to precisely reach target points in a user-interactive manner, as shown in Fig.1. To achieve this, we propose DragGAN, which consists of two main components: 1) a feature-based motion supervision that drives the handle point to move towards the target position, and 2) a new point tracking approach that leverages the discriminative generator features to keep localizing the position of the handle points. Through DragGAN, anyone can deform an image with precise control over where pixels go, thus manipulating the pose, shape, expression, and layout of diverse categories such as animals, cars, humans, landscapes, etc. As these manipulations are performed on the learned generative image manifold of a GAN, they tend to produce realistic outputs even for challenging scenarios such as hallucinating occluded content and deforming shapes that consistently follow the object's rigidity. Both qualitative and quantitative comparisons demonstrate the advantage of DragGAN over prior approaches in the tasks of image manipulation and point tracking. We also showcase the manipulation of real images through GAN inversion.
title Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2305.10973