Smart Fitting Room: A One-stop Framework for Matching-aware Virtual Try-on
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Mingzhe, Ma, Yunshan, Wu, Lei, Cheng, Kai, Li, Xue, Meng, Lei, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
by: Yu, Mingzhe, et al.
Published: (2025)
by: Yu, Mingzhe, et al.
Published: (2025)
Dual-Diffusional Generative Fashion Recommendation
by: Yu, Mingzhe, et al.
Published: (2026)
by: Yu, Mingzhe, et al.
Published: (2026)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
FashionReGen: LLM-Empowered Fashion Report Generation
by: Ding, Yujuan, et al.
Published: (2024)
by: Ding, Yujuan, et al.
Published: (2024)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
by: Chu, Meng, et al.
Published: (2023)
by: Chu, Meng, et al.
Published: (2023)
CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling
by: Ma, Yunshan, et al.
Published: (2024)
by: Ma, Yunshan, et al.
Published: (2024)
MM-Forecast: A Multimodal Approach to Temporal Event Forecasting with Large Language Models
by: Li, Haoxuan, et al.
Published: (2024)
by: Li, Haoxuan, et al.
Published: (2024)
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
by: Fang, Jianwu, et al.
Published: (2025)
by: Fang, Jianwu, et al.
Published: (2025)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
by: Wen, Haokun, et al.
Published: (2024)
by: Wen, Haokun, et al.
Published: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
by: Lan, Xing, et al.
Published: (2024)
by: Lan, Xing, et al.
Published: (2024)
Diffusion Model-Based Size Variable Virtual Try-On Technology and Evaluation Method
by: Zhang, Shufang, et al.
Published: (2025)
by: Zhang, Shufang, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
by: Liu, Zhiyuan, et al.
Published: (2023)
by: Liu, Zhiyuan, et al.
Published: (2023)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
by: Liu, Zhiyuan, et al.
Published: (2024)
by: Liu, Zhiyuan, et al.
Published: (2024)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
ReactXT: Understanding Molecular "Reaction-ship" via Reaction-Contextualized Molecule-Text Pretraining
by: Liu, Zhiyuan, et al.
Published: (2024)
by: Liu, Zhiyuan, et al.
Published: (2024)
Improving Virtual Try-On with Garment-focused Diffusion Models
by: Wan, Siqi, et al.
Published: (2024)
by: Wan, Siqi, et al.
Published: (2024)
A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning
by: Zheng, Changmeng, et al.
Published: (2024)
by: Zheng, Changmeng, et al.
Published: (2024)
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance
by: Zhu, Fengbin, et al.
Published: (2025)
by: Zhu, Fengbin, et al.
Published: (2025)
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
by: Shi, Chuancheng, et al.
Published: (2026)
by: Shi, Chuancheng, et al.
Published: (2026)
Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
by: Meng, Chunlei, et al.
Published: (2026)
by: Meng, Chunlei, et al.
Published: (2026)
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
by: Wan, Siqi, et al.
Published: (2025)
by: Wan, Siqi, et al.
Published: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Frozen LVLMs for Micro-Video Recommendation: A Systematic Study of Feature Extraction and Fusion
by: Sun, Huatuan, et al.
Published: (2025)
by: Sun, Huatuan, et al.
Published: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
by: Zhou, Yuchen, et al.
Published: (2025)
by: Zhou, Yuchen, et al.
Published: (2025)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Similar Items
-
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
by: Yu, Mingzhe, et al.
Published: (2025) -
Dual-Diffusional Generative Fashion Recommendation
by: Yu, Mingzhe, et al.
Published: (2026) -
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025) -
FashionReGen: LLM-Empowered Fashion Report Generation
by: Ding, Yujuan, et al.
Published: (2024) -
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
by: Li, Haoxuan, et al.
Published: (2025)