Key-point Guided Deformable Image Manipulation Using Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929281857224704 |
|---|---|
| author | Oh, Seok-Hwan Jung, Guil Kim, Myeong-Gee Kim, Sang-Yun Kim, Young-Min Lee, Hyeon-Jik Kwon, Hyuk-Sool Bae, Hyeon-Min |
| author_facet | Oh, Seok-Hwan Jung, Guil Kim, Myeong-Gee Kim, Sang-Yun Kim, Young-Min Lee, Hyeon-Jik Kwon, Hyuk-Sool Bae, Hyeon-Min |
| contents | In this paper, we introduce a Key-point-guided Diffusion probabilistic Model (KDM) that gains precise control over images by manipulating the object's key-point. We propose a two-stage generative model incorporating an optical flow map as an intermediate output. By doing so, a dense pixel-wise understanding of the semantic relation between the image and sparse key point is configured, leading to more realistic image generation. Additionally, the integration of optical flow helps regulate the inter-frame variance of sequential images, demonstrating an authentic sequential image generation. The KDM is evaluated with diverse key-point conditioned image synthesis tasks, including facial image generation, human pose synthesis, and echocardiography video prediction, demonstrating the KDM is proving consistency enhanced and photo-realistic images compared with state-of-the-art models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_08178 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Key-point Guided Deformable Image Manipulation Using Diffusion Model Oh, Seok-Hwan Jung, Guil Kim, Myeong-Gee Kim, Sang-Yun Kim, Young-Min Lee, Hyeon-Jik Kwon, Hyuk-Sool Bae, Hyeon-Min Computer Vision and Pattern Recognition In this paper, we introduce a Key-point-guided Diffusion probabilistic Model (KDM) that gains precise control over images by manipulating the object's key-point. We propose a two-stage generative model incorporating an optical flow map as an intermediate output. By doing so, a dense pixel-wise understanding of the semantic relation between the image and sparse key point is configured, leading to more realistic image generation. Additionally, the integration of optical flow helps regulate the inter-frame variance of sequential images, demonstrating an authentic sequential image generation. The KDM is evaluated with diverse key-point conditioned image synthesis tasks, including facial image generation, human pose synthesis, and echocardiography video prediction, demonstrating the KDM is proving consistency enhanced and photo-realistic images compared with state-of-the-art models. |
| title | Key-point Guided Deformable Image Manipulation Using Diffusion Model |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2401.08178 |