FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Rong, Li, Jinxiao, Wang, Jingnan, Zuo, Zhiwen, Dong, Jianfeng, Li, Wei, Wang, Chi, Xu, Weiwei, Wang, Xun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914244219371520
author Zhang, Rong
Li, Jinxiao
Wang, Jingnan
Zuo, Zhiwen
Dong, Jianfeng
Li, Wei
Wang, Chi
Xu, Weiwei
Wang, Xun
author_facet Zhang, Rong
Li, Jinxiao
Wang, Jingnan
Zuo, Zhiwen
Dong, Jianfeng
Li, Wei
Wang, Chi
Xu, Weiwei
Wang, Xun
contents Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
Zhang, Rong
Li, Jinxiao
Wang, Jingnan
Zuo, Zhiwen
Dong, Jianfeng
Li, Wei
Wang, Chi
Xu, Weiwei
Wang, Xun
Computer Vision and Pattern Recognition
Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods.
title FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.14031