Key-point Guided Deformable Image Manipulation Using Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oh, Seok-Hwan, Jung, Guil, Kim, Myeong-Gee, Kim, Sang-Yun, Kim, Young-Min, Lee, Hyeon-Jik, Kwon, Hyuk-Sool, Bae, Hyeon-Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929281857224704
author Oh, Seok-Hwan
Jung, Guil
Kim, Myeong-Gee
Kim, Sang-Yun
Kim, Young-Min
Lee, Hyeon-Jik
Kwon, Hyuk-Sool
Bae, Hyeon-Min
author_facet Oh, Seok-Hwan
Jung, Guil
Kim, Myeong-Gee
Kim, Sang-Yun
Kim, Young-Min
Lee, Hyeon-Jik
Kwon, Hyuk-Sool
Bae, Hyeon-Min
contents In this paper, we introduce a Key-point-guided Diffusion probabilistic Model (KDM) that gains precise control over images by manipulating the object's key-point. We propose a two-stage generative model incorporating an optical flow map as an intermediate output. By doing so, a dense pixel-wise understanding of the semantic relation between the image and sparse key point is configured, leading to more realistic image generation. Additionally, the integration of optical flow helps regulate the inter-frame variance of sequential images, demonstrating an authentic sequential image generation. The KDM is evaluated with diverse key-point conditioned image synthesis tasks, including facial image generation, human pose synthesis, and echocardiography video prediction, demonstrating the KDM is proving consistency enhanced and photo-realistic images compared with state-of-the-art models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Key-point Guided Deformable Image Manipulation Using Diffusion Model
Oh, Seok-Hwan
Jung, Guil
Kim, Myeong-Gee
Kim, Sang-Yun
Kim, Young-Min
Lee, Hyeon-Jik
Kwon, Hyuk-Sool
Bae, Hyeon-Min
Computer Vision and Pattern Recognition
In this paper, we introduce a Key-point-guided Diffusion probabilistic Model (KDM) that gains precise control over images by manipulating the object's key-point. We propose a two-stage generative model incorporating an optical flow map as an intermediate output. By doing so, a dense pixel-wise understanding of the semantic relation between the image and sparse key point is configured, leading to more realistic image generation. Additionally, the integration of optical flow helps regulate the inter-frame variance of sequential images, demonstrating an authentic sequential image generation. The KDM is evaluated with diverse key-point conditioned image synthesis tasks, including facial image generation, human pose synthesis, and echocardiography video prediction, demonstrating the KDM is proving consistency enhanced and photo-realistic images compared with state-of-the-art models.
title Key-point Guided Deformable Image Manipulation Using Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.08178