VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Jun, Zhao, Fuwei, Xu, Youjiang, Dong, Xin, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
by: Zheng, Jun, et al.
Published: (2024)
by: Zheng, Jun, et al.
Published: (2024)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
by: Wan, Zhenchen, et al.
Published: (2024)
by: Wan, Zhenchen, et al.
Published: (2024)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)
by: Lyu, Hengye, et al.
Published: (2026)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
by: Zuo, Tongchun, et al.
Published: (2025)
by: Zuo, Tongchun, et al.
Published: (2025)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
by: Yang, Lianwei, et al.
Published: (2025)
by: Yang, Lianwei, et al.
Published: (2025)
DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos
by: Shen, Wenhao, et al.
Published: (2026)
by: Shen, Wenhao, et al.
Published: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild
by: Zhou, Jingkai, et al.
Published: (2025)
by: Zhou, Jingkai, et al.
Published: (2025)
Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection
by: Pan, Yanjie, et al.
Published: (2025)
by: Pan, Yanjie, et al.
Published: (2025)
DiVE: DiT-based Video Generation with Enhanced Control
by: Jiang, Junpeng, et al.
Published: (2024)
by: Jiang, Junpeng, et al.
Published: (2024)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
by: Qi, Tianhao, et al.
Published: (2025)
by: Qi, Tianhao, et al.
Published: (2025)
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
by: Zheng, Jun, et al.
Published: (2026)
by: Zheng, Jun, et al.
Published: (2026)
DiT-VTON: Diffusion Transformer Framework for Unified Multi-Category Virtual Try-On and Virtual Try-All with Integrated Image Editing
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
by: Duan, Zheng-Peng, et al.
Published: (2025)
by: Duan, Zheng-Peng, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
EfficientVITON: An Efficient Virtual Try-On Model using Optimized Diffusion Process
by: Atef, Mostafa, et al.
Published: (2025)
by: Atef, Mostafa, et al.
Published: (2025)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
by: Zou, Cheng, et al.
Published: (2025)
by: Zou, Cheng, et al.
Published: (2025)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
$Δ$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2024)
by: Chen, Pengtao, et al.
Published: (2024)
VITON-DRR: Details Retention Virtual Try-on via Non-rigid Registration
by: Li, Ben, et al.
Published: (2025)
by: Li, Ben, et al.
Published: (2025)
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
by: Lin, Ente, et al.
Published: (2024)
by: Lin, Ente, et al.
Published: (2024)
ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
by: Park, Junseo, et al.
Published: (2025)
by: Park, Junseo, et al.
Published: (2025)
EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
by: Zhu, Rui, et al.
Published: (2024)
by: Zhu, Rui, et al.
Published: (2024)
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers
by: Chen, Ruichen, et al.
Published: (2025)
by: Chen, Ruichen, et al.
Published: (2025)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
Similar Items
-
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
by: Zheng, Jun, et al.
Published: (2024) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024) -
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026) -
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
by: Wan, Zhenchen, et al.
Published: (2024) -
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)