CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chong, Zheng, Dong, Xiao, Li, Haoxiang, Zhang, Shiyue, Zhang, Wenqing, Zhang, Xujie, Zhao, Hanqing, Jiang, Dongmei, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
von: Chong, Zheng, et al.
Veröffentlicht: (2025)
DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images
von: Wang, Hualiang, et al.
Veröffentlicht: (2024)
von: Wang, Hualiang, et al.
Veröffentlicht: (2024)
Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation
von: Egorov, Konstantin, et al.
Veröffentlicht: (2025)
von: Egorov, Konstantin, et al.
Veröffentlicht: (2025)
Efficient Vision-based Vehicle Speed Estimation
von: Macko, Andrej, et al.
Veröffentlicht: (2025)
von: Macko, Andrej, et al.
Veröffentlicht: (2025)
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
von: Fang, Qihang, et al.
Veröffentlicht: (2024)
von: Fang, Qihang, et al.
Veröffentlicht: (2024)
PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
von: Hou, Yang, et al.
Veröffentlicht: (2024)
von: Hou, Yang, et al.
Veröffentlicht: (2024)
Stylized Face Sketch Extraction via Generative Prior with Limited Data
von: Yun, Kwan, et al.
Veröffentlicht: (2024)
von: Yun, Kwan, et al.
Veröffentlicht: (2024)
SymFace: Additional Facial Symmetry Loss for Deep Face Recognition
von: Prakash, Pritesh, et al.
Veröffentlicht: (2024)
von: Prakash, Pritesh, et al.
Veröffentlicht: (2024)
LeGO: Leveraging a Surface Deformation Network for Animatable Stylized Face Generation with One Example
von: Yoon, Soyeon, et al.
Veröffentlicht: (2024)
von: Yoon, Soyeon, et al.
Veröffentlicht: (2024)
Distilling foundation models for robust and efficient models in digital pathology
von: Filiot, Alexandre, et al.
Veröffentlicht: (2025)
von: Filiot, Alexandre, et al.
Veröffentlicht: (2025)
Harnessing Deep Learning and Satellite Imagery for Post-Buyout Land Cover Mapping
von: Otal, Hakan T., et al.
Veröffentlicht: (2024)
von: Otal, Hakan T., et al.
Veröffentlicht: (2024)
Detailed Evaluation of Modern Machine Learning Approaches for Optic Plastics Sorting
von: Maheshkar, Vaishali, et al.
Veröffentlicht: (2025)
von: Maheshkar, Vaishali, et al.
Veröffentlicht: (2025)
Empowering Image Recovery_ A Multi-Attention Approach
von: Wen, Juan, et al.
Veröffentlicht: (2024)
von: Wen, Juan, et al.
Veröffentlicht: (2024)
SelectiveKD: A semi-supervised framework for cancer detection in DBT through Knowledge Distillation and Pseudo-labeling
von: Dillard, Laurent, et al.
Veröffentlicht: (2024)
von: Dillard, Laurent, et al.
Veröffentlicht: (2024)
DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
von: Çalışkan, Halil Hüseyin, et al.
Veröffentlicht: (2025)
von: Çalışkan, Halil Hüseyin, et al.
Veröffentlicht: (2025)
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
von: Choi, Seonghwa, et al.
Veröffentlicht: (2025)
von: Choi, Seonghwa, et al.
Veröffentlicht: (2025)
Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
von: Khan, Md Ashik, et al.
Veröffentlicht: (2025)
von: Khan, Md Ashik, et al.
Veröffentlicht: (2025)
VitalLens 2.0: High-Fidelity rPPG for Heart Rate Variability Estimation from Face Video
von: Rouast, Philipp V.
Veröffentlicht: (2025)
von: Rouast, Philipp V.
Veröffentlicht: (2025)
N-DriverMotion: Driver motion learning and prediction using an event-based camera and directly trained spiking neural networks on Loihi 2
von: Chung, Hyo Jong, et al.
Veröffentlicht: (2024)
von: Chung, Hyo Jong, et al.
Veröffentlicht: (2024)
Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives
von: He, Shijing, et al.
Veröffentlicht: (2025)
von: He, Shijing, et al.
Veröffentlicht: (2025)
Event-ECC: Asynchronous Tracking of Events with Continuous Optimization
von: Zafeiri, Maria, et al.
Veröffentlicht: (2024)
von: Zafeiri, Maria, et al.
Veröffentlicht: (2024)
Knowledge Distillation: Enhancing Neural Network Compression with Integrated Gradients
von: Hernandez, David E., et al.
Veröffentlicht: (2025)
von: Hernandez, David E., et al.
Veröffentlicht: (2025)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
Geochemistry of volcanic rocks from the Hikurangi and Manihiki Plateaus
von: Hoernle, Kaj, et al.
Veröffentlicht: (2010)
von: Hoernle, Kaj, et al.
Veröffentlicht: (2010)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
von: Ömercikoğlu, Ahmet Can, et al.
Veröffentlicht: (2025)
von: Ömercikoğlu, Ahmet Can, et al.
Veröffentlicht: (2025)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
von: Long, Yuchong, et al.
Veröffentlicht: (2025)
von: Long, Yuchong, et al.
Veröffentlicht: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
Probabilistic Selection in AgentSpeak(L)
von: Coelho, Francisco, et al.
Veröffentlicht: (2014)
von: Coelho, Francisco, et al.
Veröffentlicht: (2014)
Eleven Primitives and Three Gates: The Universal Structure of Computational Imaging
von: Yang, Chengshuai, et al.
Veröffentlicht: (2026)
von: Yang, Chengshuai, et al.
Veröffentlicht: (2026)
Model compression using knowledge distillation with integrated gradients
von: Hernandez, David E., et al.
Veröffentlicht: (2025)
von: Hernandez, David E., et al.
Veröffentlicht: (2025)
Balanced conic rectified flow
von: Kim, Shin Seong, et al.
Veröffentlicht: (2025)
von: Kim, Shin Seong, et al.
Veröffentlicht: (2025)
Motion-Based Sign Language Video Summarization using Curvature and Torsion
von: Sartinas, Evangelos G., et al.
Veröffentlicht: (2023)
von: Sartinas, Evangelos G., et al.
Veröffentlicht: (2023)
BG-YOLO: A Bidirectional-Guided Method for Underwater Object Detection
von: Zhang, Jian, et al.
Veröffentlicht: (2024)
von: Zhang, Jian, et al.
Veröffentlicht: (2024)
A Plug-and-Play Method with Inpainting Network for Bayesian Uncertainty Quantification in Imaging
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2023)
EditP23: 3D Editing via Propagation of Image Prompts to Multi-View
von: Bar-On, Roi, et al.
Veröffentlicht: (2025)
von: Bar-On, Roi, et al.
Veröffentlicht: (2025)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
von: Merugu, Ranjith, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
von: Chong, Zheng, et al.
Veröffentlicht: (2025) -
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
von: Chong, Zheng, et al.
Veröffentlicht: (2025) -
DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
von: Zhang, Rui, et al.
Veröffentlicht: (2025) -
Tri-Plane Mamba: Efficiently Adapting Segment Anything Model for 3D Medical Images
von: Wang, Hualiang, et al.
Veröffentlicht: (2024) -
Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation
von: Egorov, Konstantin, et al.
Veröffentlicht: (2025)