LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Girella, Federico, Talon, Davide, Liu, Ziyue, Ruan, Zanxi, Wang, Yiming, Cristani, Marco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908518846562304
author Girella, Federico
Talon, Davide
Liu, Ziyue
Ruan, Zanxi
Wang, Yiming
Cristani, Marco
author_facet Girella, Federico
Talon, Davide
Liu, Ziyue
Ruan, Zanxi
Wang, Yiming
Cristani, Marco
contents Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing material, texture, and stylistic details. In this paper, we present LOcalized Text and Sketch for fashion image generation (LOTS), an approach for compositional sketch-text based generation of complete fashion outlooks. LOTS leverages a global description with paired localized sketch + text information for conditioning and introduces a novel step-based merging strategy for diffusion adaptation. First, a Modularized Pair-Centric representation encodes sketches and text into a shared latent space while preserving independent localized features; then, a Diffusion Pair Guidance phase integrates both local and global conditioning via attention-based guidance within the diffusion model's multi-step denoising process. To validate our method, we build on Fashionpedia to release Sketchy, the first fashion dataset where multiple text-sketch pairs are provided per image. Quantitative results show LOTS achieves state-of-the-art image generation performance on both global and localized metrics, while qualitative examples and a human evaluation study highlight its unprecedented level of design customization.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
Girella, Federico
Talon, Davide
Liu, Ziyue
Ruan, Zanxi
Wang, Yiming
Cristani, Marco
Computer Vision and Pattern Recognition
Artificial Intelligence
Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing material, texture, and stylistic details. In this paper, we present LOcalized Text and Sketch for fashion image generation (LOTS), an approach for compositional sketch-text based generation of complete fashion outlooks. LOTS leverages a global description with paired localized sketch + text information for conditioning and introduces a novel step-based merging strategy for diffusion adaptation. First, a Modularized Pair-Centric representation encodes sketches and text into a shared latent space while preserving independent localized features; then, a Diffusion Pair Guidance phase integrates both local and global conditioning via attention-based guidance within the diffusion model's multi-step denoising process. To validate our method, we build on Fashionpedia to release Sketchy, the first fashion dataset where multiple text-sketch pairs are provided per image. Quantitative results show LOTS achieves state-of-the-art image generation performance on both global and localized metrics, while qualitative examples and a human evaluation study highlight its unprecedented level of design customization.
title LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.22627