DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lin, Ente, Zhang, Xujie, Zhao, Fuwei, Luo, Yuxuan, Dong, Xin, Zeng, Long, Liang, Xiaodan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929681490509824
author Lin, Ente
Zhang, Xujie
Zhao, Fuwei
Luo, Yuxuan
Dong, Xin
Zeng, Long
Liang, Xiaodan
author_facet Lin, Ente
Zhang, Xujie
Zhao, Fuwei
Luo, Yuxuan
Dong, Xin
Zeng, Long
Liang, Xiaodan
contents Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) \textbf{Lightweight training}: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2)\textbf{Anything-Dressing}: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) \textbf{Plug-and-play}: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both $768 \times 512$ high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17644
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
Lin, Ente
Zhang, Xujie
Zhao, Fuwei
Luo, Yuxuan
Dong, Xin
Zeng, Long
Liang, Xiaodan
Computer Vision and Pattern Recognition
Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) \textbf{Lightweight training}: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2)\textbf{Anything-Dressing}: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) \textbf{Plug-and-play}: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both $768 \times 512$ high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation.
title DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.17644