AIpparel: A Multimodal Foundation Model for Digital Garments

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nakayama, Kiyohiro, Ackermann, Jan, Kesdogan, Timur Levent, Zheng, Yang, Korosteleva, Maria, Sorkine-Hornung, Olga, Guibas, Leonidas J., Yang, Guandao, Wetzstein, Gordon
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910903943823360
author Nakayama, Kiyohiro
Ackermann, Jan
Kesdogan, Timur Levent
Zheng, Yang
Korosteleva, Maria
Sorkine-Hornung, Olga
Guibas, Leonidas J.
Yang, Guandao
Wetzstein, Gordon
author_facet Nakayama, Kiyohiro
Ackermann, Jan
Kesdogan, Timur Levent
Zheng, Yang
Korosteleva, Maria
Sorkine-Hornung, Olga
Guibas, Leonidas J.
Yang, Guandao
Wetzstein, Gordon
contents Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming process, largely due to the manual work involved in designing them. To simplify this process, we introduce AIpparel, a multimodal foundation model for generating and editing sewing patterns. Our model fine-tunes state-of-the-art large multimodal models (LMMs) on a custom-curated large-scale dataset of over 120,000 unique garments, each with multimodal annotations including text, images, and sewing patterns. Additionally, we propose a novel tokenization scheme that concisely encodes these complex sewing patterns so that LLMs can learn to predict them efficiently. AIpparel achieves state-of-the-art performance in single-modal tasks, including text-to-garment and image-to-garment prediction, and enables novel multimodal garment generation applications such as interactive garment editing. The project website is at https://georgenakayama.github.io/AIpparel/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AIpparel: A Multimodal Foundation Model for Digital Garments
Nakayama, Kiyohiro
Ackermann, Jan
Kesdogan, Timur Levent
Zheng, Yang
Korosteleva, Maria
Sorkine-Hornung, Olga
Guibas, Leonidas J.
Yang, Guandao
Wetzstein, Gordon
Computer Vision and Pattern Recognition
Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming process, largely due to the manual work involved in designing them. To simplify this process, we introduce AIpparel, a multimodal foundation model for generating and editing sewing patterns. Our model fine-tunes state-of-the-art large multimodal models (LMMs) on a custom-curated large-scale dataset of over 120,000 unique garments, each with multimodal annotations including text, images, and sewing patterns. Additionally, we propose a novel tokenization scheme that concisely encodes these complex sewing patterns so that LLMs can learn to predict them efficiently. AIpparel achieves state-of-the-art performance in single-modal tasks, including text-to-garment and image-to-garment prediction, and enables novel multimodal garment generation applications such as interactive garment editing. The project website is at https://georgenakayama.github.io/AIpparel/.
title AIpparel: A Multimodal Foundation Model for Digital Garments
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.03937