Scene Aware Person Image Generation through Global Contextual Conditioning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Roy, Prasun, Ghosh, Subhankar, Bhattacharya, Saumik, Pal, Umapada, Blumenstein, Michael
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917926212206592
author Roy, Prasun
Ghosh, Subhankar
Bhattacharya, Saumik
Pal, Umapada
Blumenstein, Michael
author_facet Roy, Prasun
Ghosh, Subhankar
Bhattacharya, Saumik
Pal, Umapada
Blumenstein, Michael
contents Person image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results.
format Preprint
id arxiv_https___arxiv_org_abs_2206_02717
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Scene Aware Person Image Generation through Global Contextual Conditioning
Roy, Prasun
Ghosh, Subhankar
Bhattacharya, Saumik
Pal, Umapada
Blumenstein, Michael
Computer Vision and Pattern Recognition
Multimedia
Person image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results.
title Scene Aware Person Image Generation through Global Contextual Conditioning
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2206.02717