Saved in:
| Main Authors: | Tao, Ming, Bao, Bing-Kun, Wang, Yaowei, Xu, Changsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.05619 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
by: Tao, Ming, et al.
Published: (2024)
by: Tao, Ming, et al.
Published: (2024)
Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
by: Xu, Mengling, et al.
Published: (2025)
by: Xu, Mengling, et al.
Published: (2025)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)
by: Shi, Baifeng, et al.
Published: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?
by: Yan, Renye, et al.
Published: (2026)
by: Yan, Renye, et al.
Published: (2026)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
ITVTON: Virtual Try-On Diffusion Transformer Based on Integrated Image and Text
by: Ni, Haifeng, et al.
Published: (2025)
by: Ni, Haifeng, et al.
Published: (2025)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
FLDM-VTON: Faithful Latent Diffusion Model for Virtual Try-on
by: Wang, Chenhui, et al.
Published: (2024)
by: Wang, Chenhui, et al.
Published: (2024)
Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
by: Zou, Cheng, et al.
Published: (2025)
by: Zou, Cheng, et al.
Published: (2025)
OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on
by: Xu, Yuhao, et al.
Published: (2024)
by: Xu, Yuhao, et al.
Published: (2024)
Do Satellite Tasks Need Special Pretraining?
by: Vanyan, Ani, et al.
Published: (2025)
by: Vanyan, Ani, et al.
Published: (2025)
Do We Need Reformer for Vision? An Experimental Comparison with Vision Transformers
by: Bellaj, Ali El, et al.
Published: (2025)
by: Bellaj, Ali El, et al.
Published: (2025)
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
by: Wan, Zhenchen, et al.
Published: (2024)
by: Wan, Zhenchen, et al.
Published: (2024)
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
by: Wan, Siqi, et al.
Published: (2025)
by: Wan, Siqi, et al.
Published: (2025)
How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models
by: Hackel, Leonard, et al.
Published: (2026)
by: Hackel, Leonard, et al.
Published: (2026)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Pilot: Building the Federated Multimodal Instruction Tuning Framework
by: Xiong, Baochen, et al.
Published: (2025)
by: Xiong, Baochen, et al.
Published: (2025)
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Improving Virtual Try-On with Garment-focused Diffusion Models
by: Wan, Siqi, et al.
Published: (2024)
by: Wan, Siqi, et al.
Published: (2024)
Do We Need Perfect Data? Leveraging Noise for Domain Generalized Segmentation
by: Kim, Taeyeong, et al.
Published: (2025)
by: Kim, Taeyeong, et al.
Published: (2025)
Motion-aware Latent Diffusion Models for Video Frame Interpolation
by: Huang, Zhilin, et al.
Published: (2024)
by: Huang, Zhilin, et al.
Published: (2024)
Can We Predict Performance of Large Models across Vision-Language Tasks?
by: Zhao, Qinyu, et al.
Published: (2024)
by: Zhao, Qinyu, et al.
Published: (2024)
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
by: Tao, Xijia, et al.
Published: (2024)
by: Tao, Xijia, et al.
Published: (2024)
Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting
by: Guo, Mingyue, et al.
Published: (2024)
by: Guo, Mingyue, et al.
Published: (2024)
Pixel Motion Diffusion is What We Need for Robot Control
by: Nguyen, E-Ro, et al.
Published: (2025)
by: Nguyen, E-Ro, et al.
Published: (2025)
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
by: Wang, Zhendong, et al.
Published: (2025)
by: Wang, Zhendong, et al.
Published: (2025)
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
by: Zhao, Zhonghan, et al.
Published: (2024)
by: Zhao, Zhonghan, et al.
Published: (2024)
Do We Really Need a Large Number of Visual Prompts?
by: Kim, Youngeun, et al.
Published: (2023)
by: Kim, Youngeun, et al.
Published: (2023)
CAT-DM: Controllable Accelerated Virtual Try-on with Diffusion Model
by: Zeng, Jianhao, et al.
Published: (2023)
by: Zeng, Jianhao, et al.
Published: (2023)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025)
by: Liu, Jinxi, et al.
Published: (2025)
Struggle with Adversarial Defense? Try Diffusion
by: Li, Yujie, et al.
Published: (2024)
by: Li, Yujie, et al.
Published: (2024)
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
by: Guo, Hanzhong, et al.
Published: (2024)
by: Guo, Hanzhong, et al.
Published: (2024)
Similar Items
-
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
by: Tao, Ming, et al.
Published: (2024) -
Chain-of-Cooking:Cooking Process Visualization via Bidirectional Chain-of-Thought Guidance
by: Xu, Mengling, et al.
Published: (2025) -
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
by: Zou, Jiayi, et al.
Published: (2025) -
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024) -
When Do We Not Need Larger Vision Models?
by: Shi, Baifeng, et al.
Published: (2024)