One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916394385276928 |
|---|---|
| author | Fan, Dongqi Chen, Tao Wang, Mingjie Ma, Rui Tang, Qiang Yi, Zili Wang, Qian Chang, Liang |
| author_facet | Fan, Dongqi Chen, Tao Wang, Mingjie Ma, Rui Tang, Qiang Yi, Zili Wang, Qian Chang, Liang |
| contents | Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_09593 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild Fan, Dongqi Chen, Tao Wang, Mingjie Ma, Rui Tang, Qiang Yi, Zili Wang, Qian Chang, Liang Computer Vision and Pattern Recognition Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU. |
| title | One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.09593 |