Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cai, Pengxiang, Liu, Zhiwei, Zhu, Guibo, Niu, Yunfang, Wang, Jinqiao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910543680372736
author Cai, Pengxiang
Liu, Zhiwei
Zhu, Guibo
Niu, Yunfang
Wang, Jinqiao
author_facet Cai, Pengxiang
Liu, Zhiwei
Zhu, Guibo
Niu, Yunfang
Wang, Jinqiao
contents Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fail to achieve pixel-level fine-grained control, or their inference speed requires optimization. To address this, this paper for the first time employs a regression-based network to learn the variation patterns of StyleGAN latent codes during the image dragging process. This method enables pixel-level precision in dragging editing with little time cost. Users can specify handle points and their corresponding target points on any GAN-generated images, and our method will move each handle point to its corresponding target point. Through experimental analysis, we discover that a short movement distance from handle points to target points yields a high-fidelity edited image, as the model only needs to predict the movement of a small portion of pixels. To achieve this, we decompose the entire movement process into multiple sub-processes. Specifically, we develop a transformer encoder-decoder based network named 'Latent Predictor' to predict the latent code motion trajectories from handle points to target points in an autoregressive manner. Moreover, to enhance the prediction stability, we introduce a component named 'Latent Regularizer', aimed at constraining the latent code motion within the distribution of natural images. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) inference speed and image editing performance at the pixel-level granularity.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18656
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
Cai, Pengxiang
Liu, Zhiwei
Zhu, Guibo
Niu, Yunfang
Wang, Jinqiao
Computer Vision and Pattern Recognition
Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fail to achieve pixel-level fine-grained control, or their inference speed requires optimization. To address this, this paper for the first time employs a regression-based network to learn the variation patterns of StyleGAN latent codes during the image dragging process. This method enables pixel-level precision in dragging editing with little time cost. Users can specify handle points and their corresponding target points on any GAN-generated images, and our method will move each handle point to its corresponding target point. Through experimental analysis, we discover that a short movement distance from handle points to target points yields a high-fidelity edited image, as the model only needs to predict the movement of a small portion of pixels. To achieve this, we decompose the entire movement process into multiple sub-processes. Specifically, we develop a transformer encoder-decoder based network named 'Latent Predictor' to predict the latent code motion trajectories from handle points to target points in an autoregressive manner. Moreover, to enhance the prediction stability, we introduce a component named 'Latent Regularizer', aimed at constraining the latent code motion within the distribution of natural images. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) inference speed and image editing performance at the pixel-level granularity.
title Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18656