Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912718527660032 |
|---|---|
| author | Xie, Chengyu Gong, Zhi Ren, Junchi Yu, Linkun Shen, Si Shen, Fei Du, Xiaoyu |
| author_facet | Xie, Chengyu Gong, Zhi Ren, Junchi Yu, Linkun Shen, Si Shen, Fei Du, Xiaoyu |
| contents | Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior module (APM) infers a holistic identity preserving prior from incomplete references, and the joint conditional injection (JCI) mechanism fuses multi-view cues and injects shared conditioning into the denoising backbone to align identity, color, and texture across poses. JCDM supports a variable number of reference views and integrates with standard diffusion backbones with minimal and targeted architectural modifications. Experiments demonstrate state of the art fidelity and cross-view consistency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_15092 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis Xie, Chengyu Gong, Zhi Ren, Junchi Yu, Linkun Shen, Si Shen, Fei Du, Xiaoyu Computer Vision and Pattern Recognition Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior module (APM) infers a holistic identity preserving prior from incomplete references, and the joint conditional injection (JCI) mechanism fuses multi-view cues and injects shared conditioning into the denoising backbone to align identity, color, and texture across poses. JCDM supports a variable number of reference views and integrates with standard diffusion backbones with minimal and targeted architectural modifications. Experiments demonstrate state of the art fidelity and cross-view consistency. |
| title | Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.15092 |