One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Dongqi, Chen, Tao, Wang, Mingjie, Ma, Rui, Tang, Qiang, Yi, Zili, Wang, Qian, Chang, Liang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916394385276928
author Fan, Dongqi
Chen, Tao
Wang, Mingjie
Ma, Rui
Tang, Qiang
Yi, Zili
Wang, Qian
Chang, Liang
author_facet Fan, Dongqi
Chen, Tao
Wang, Mingjie
Ma, Rui
Tang, Qiang
Yi, Zili
Wang, Qian
Chang, Liang
contents Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09593
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
Fan, Dongqi
Chen, Tao
Wang, Mingjie
Ma, Rui
Tang, Qiang
Yi, Zili
Wang, Qian
Chang, Liang
Computer Vision and Pattern Recognition
Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU.
title One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.09593