Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Yu, Zhu, Ting, Liu, Yichun, Ma, Lichen, Shan, Xinyuan, Fu, Jingling, Shi, Yu, Huang, Junshi, Li, Yan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918499361751040
author He, Yu
Zhu, Ting
Liu, Yichun
Ma, Lichen
Shan, Xinyuan
Fu, Jingling
Shi, Yu
Huang, Junshi
Li, Yan
author_facet He, Yu
Zhu, Ting
Liu, Yichun
Ma, Lichen
Shan, Xinyuan
Fu, Jingling
Shi, Yu
Huang, Junshi
Li, Yan
contents Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored, requiring comprehensive e-commercial dataset and elaborative utilization of multi-modal condition. In this paper, we propose a brand-new e-commerce dataset, named Fashion130k, with various occasions, models, and garment types. For the consistent generation of garment, we design a framework with Unified Multi-modal Condition (UMC) to align and integrate the text and visual prompts into generation model. Specifically, we explore an embedding refiner to extract the unified embeddings of multi-modal prompts, within which a Fusion Transformer is proposed to align the multi-modal embeddings by adjusting the modality gap between text and image. Based on unified embeddings, the attention in generation model is redesigned to emphasis the correlations between prompts and noise image, inducing that the noise image can select the pivotal tokens of prompts for consistent outfit generation. Our dataset and proposed framework offer a general and nuanced exploration of multi-modal prompts for generation models. Extensive experiments on real-world applications and benchmark demonstrate the effectiveness of UMC in visual consistency, achieving promising result than that of SoTA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10127
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
He, Yu
Zhu, Ting
Liu, Yichun
Ma, Lichen
Shan, Xinyuan
Fu, Jingling
Shi, Yu
Huang, Junshi
Li, Yan
Computer Vision and Pattern Recognition
Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image and text prompt. However, the potential of outfit generation remains underexplored, requiring comprehensive e-commercial dataset and elaborative utilization of multi-modal condition. In this paper, we propose a brand-new e-commerce dataset, named Fashion130k, with various occasions, models, and garment types. For the consistent generation of garment, we design a framework with Unified Multi-modal Condition (UMC) to align and integrate the text and visual prompts into generation model. Specifically, we explore an embedding refiner to extract the unified embeddings of multi-modal prompts, within which a Fusion Transformer is proposed to align the multi-modal embeddings by adjusting the modality gap between text and image. Based on unified embeddings, the attention in generation model is redesigned to emphasis the correlations between prompts and noise image, inducing that the noise image can select the pivotal tokens of prompts for consistent outfit generation. Our dataset and proposed framework offer a general and nuanced exploration of multi-modal prompts for generation models. Extensive experiments on real-world applications and benchmark demonstrate the effectiveness of UMC in visual consistency, achieving promising result than that of SoTA methods.
title Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10127