Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Chengyu, Gong, Zhi, Ren, Junchi, Yu, Linkun, Shen, Si, Shen, Fei, Du, Xiaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912718527660032
author Xie, Chengyu
Gong, Zhi
Ren, Junchi
Yu, Linkun
Shen, Si
Shen, Fei
Du, Xiaoyu
author_facet Xie, Chengyu
Gong, Zhi
Ren, Junchi
Yu, Linkun
Shen, Si
Shen, Fei
Du, Xiaoyu
contents Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior module (APM) infers a holistic identity preserving prior from incomplete references, and the joint conditional injection (JCI) mechanism fuses multi-view cues and injects shared conditioning into the denoising backbone to align identity, color, and texture across poses. JCDM supports a variable number of reference views and integrates with standard diffusion backbones with minimal and targeted architectural modifications. Experiments demonstrate state of the art fidelity and cross-view consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15092
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis
Xie, Chengyu
Gong, Zhi
Ren, Junchi
Yu, Linkun
Shen, Si
Shen, Fei
Du, Xiaoyu
Computer Vision and Pattern Recognition
Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior module (APM) infers a holistic identity preserving prior from incomplete references, and the joint conditional injection (JCI) mechanism fuses multi-view cues and injects shared conditioning into the denoising backbone to align identity, color, and texture across poses. JCDM supports a variable number of reference views and integrates with standard diffusion backbones with minimal and targeted architectural modifications. Experiments demonstrate state of the art fidelity and cross-view consistency.
title Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15092