Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cha, Hyunsoo, Woo, Wonjung, Kim, Byungjun, Joo, Hanbyul
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917456172285952
author Cha, Hyunsoo
Woo, Wonjung
Kim, Byungjun
Joo, Hanbyul
author_facet Cha, Hyunsoo
Woo, Wonjung
Kim, Byungjun
Joo, Hanbyul
contents We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two-stage pipelines treat image-based virtual try-on and pose-driven animation as separate processes, which often results in identity drift, garment distortion, and front-back inconsistency. Our model addresses these issues by performing the entire process in a single unified step to achieve coherent synthesis. To enable this setting, we construct large-scale triplet supervision. Our data generation pipeline includes generating identity-preserving human images in alternative outfits that differ from garment catalog images, capturing full upper and lower garment triplets to overcome the single-garment-posed video pair limitation, and assembling diverse in-the-wild triplets without requiring garment catalog images. We further introduce a Dual Module architecture for video diffusion transformers to stabilize training, preserve pretrained generative quality, and improve garment accuracy, pose adherence, and identity preservation while supporting zero-shot garment interpolation. Together, these contributions allow Vanast to produce high-fidelity, identity-consistent animation across a wide range of garment types.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04934
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
Cha, Hyunsoo
Woo, Wonjung
Kim, Byungjun
Joo, Hanbyul
Computer Vision and Pattern Recognition
We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two-stage pipelines treat image-based virtual try-on and pose-driven animation as separate processes, which often results in identity drift, garment distortion, and front-back inconsistency. Our model addresses these issues by performing the entire process in a single unified step to achieve coherent synthesis. To enable this setting, we construct large-scale triplet supervision. Our data generation pipeline includes generating identity-preserving human images in alternative outfits that differ from garment catalog images, capturing full upper and lower garment triplets to overcome the single-garment-posed video pair limitation, and assembling diverse in-the-wild triplets without requiring garment catalog images. We further introduce a Dual Module architecture for video diffusion transformers to stabilize training, preserve pretrained generative quality, and improve garment accuracy, pose adherence, and identity preservation while supporting zero-shot garment interpolation. Together, these contributions allow Vanast to produce high-fidelity, identity-consistent animation across a wide range of garment types.
title Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04934