AIpparel: A Multimodal Foundation Model for Digital Garments
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910903943823360 |
|---|---|
| author | Nakayama, Kiyohiro Ackermann, Jan Kesdogan, Timur Levent Zheng, Yang Korosteleva, Maria Sorkine-Hornung, Olga Guibas, Leonidas J. Yang, Guandao Wetzstein, Gordon |
| author_facet | Nakayama, Kiyohiro Ackermann, Jan Kesdogan, Timur Levent Zheng, Yang Korosteleva, Maria Sorkine-Hornung, Olga Guibas, Leonidas J. Yang, Guandao Wetzstein, Gordon |
| contents | Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming process, largely due to the manual work involved in designing them. To simplify this process, we introduce AIpparel, a multimodal foundation model for generating and editing sewing patterns. Our model fine-tunes state-of-the-art large multimodal models (LMMs) on a custom-curated large-scale dataset of over 120,000 unique garments, each with multimodal annotations including text, images, and sewing patterns. Additionally, we propose a novel tokenization scheme that concisely encodes these complex sewing patterns so that LLMs can learn to predict them efficiently. AIpparel achieves state-of-the-art performance in single-modal tasks, including text-to-garment and image-to-garment prediction, and enables novel multimodal garment generation applications such as interactive garment editing. The project website is at https://georgenakayama.github.io/AIpparel/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_03937 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AIpparel: A Multimodal Foundation Model for Digital Garments Nakayama, Kiyohiro Ackermann, Jan Kesdogan, Timur Levent Zheng, Yang Korosteleva, Maria Sorkine-Hornung, Olga Guibas, Leonidas J. Yang, Guandao Wetzstein, Gordon Computer Vision and Pattern Recognition Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming process, largely due to the manual work involved in designing them. To simplify this process, we introduce AIpparel, a multimodal foundation model for generating and editing sewing patterns. Our model fine-tunes state-of-the-art large multimodal models (LMMs) on a custom-curated large-scale dataset of over 120,000 unique garments, each with multimodal annotations including text, images, and sewing patterns. Additionally, we propose a novel tokenization scheme that concisely encodes these complex sewing patterns so that LLMs can learn to predict them efficiently. AIpparel achieves state-of-the-art performance in single-modal tasks, including text-to-garment and image-to-garment prediction, and enables novel multimodal garment generation applications such as interactive garment editing. The project website is at https://georgenakayama.github.io/AIpparel/. |
| title | AIpparel: A Multimodal Foundation Model for Digital Garments |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.03937 |