Saved in:
Bibliographic Details
Main Authors: Yin, Deqiang, Guo, Junyi, Lu, Huanda, Wu, Fangyu, Lu, Dongming
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.03497
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911105210646528
author Yin, Deqiang
Guo, Junyi
Lu, Huanda
Wu, Fangyu
Lu, Dongming
author_facet Yin, Deqiang
Guo, Junyi
Lu, Huanda
Wu, Fangyu
Lu, Dongming
contents Instruction-based garment editing enables precise image modifications via natural language, with broad applications in fashion design and customization. Unlike general editing tasks, it requires understanding garment-specific semantics and attribute dependencies. However, progress is limited by the scarcity of high-quality instruction-image pairs, as manual annotation is costly and hard to scale. While MLLMs have shown promise in automated data synthesis, their application to garment editing is constrained by imprecise instruction modeling and a lack of fashion-specific supervisory signals. To address these challenges, we present an automated pipeline for constructing a garment editing dataset. We first define six editing instruction categories aligned with real-world fashion workflows to guide the generation of balanced and diverse instruction-image triplets. Second, we introduce Fashion Edit Score, a semantic-aware evaluation metric that captures semantic dependencies between garment attributes and provides reliable supervision during construction. Using this pipeline, we construct a total of 52,257 candidate triplets and retain 20,596 high-quality triplets to build EditGarment, the first instruction-based dataset tailored to standalone garment editing. The project page is https://yindq99.github.io/EditGarment-project/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware Evaluation
Yin, Deqiang
Guo, Junyi
Lu, Huanda
Wu, Fangyu
Lu, Dongming
Computer Vision and Pattern Recognition
Instruction-based garment editing enables precise image modifications via natural language, with broad applications in fashion design and customization. Unlike general editing tasks, it requires understanding garment-specific semantics and attribute dependencies. However, progress is limited by the scarcity of high-quality instruction-image pairs, as manual annotation is costly and hard to scale. While MLLMs have shown promise in automated data synthesis, their application to garment editing is constrained by imprecise instruction modeling and a lack of fashion-specific supervisory signals. To address these challenges, we present an automated pipeline for constructing a garment editing dataset. We first define six editing instruction categories aligned with real-world fashion workflows to guide the generation of balanced and diverse instruction-image triplets. Second, we introduce Fashion Edit Score, a semantic-aware evaluation metric that captures semantic dependencies between garment attributes and provides reliable supervision during construction. Using this pipeline, we construct a total of 52,257 candidate triplets and retain 20,596 high-quality triplets to build EditGarment, the first instruction-based dataset tailored to standalone garment editing. The project page is https://yindq99.github.io/EditGarment-project/.
title EditGarment: An Instruction-Based Garment Editing Dataset Constructed with Automated MLLM Synthesis and Semantic-Aware Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03497