Type Aware Embedding for Outfit Compatibility with CLIP

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Dae Lim, Chung
Format: Recurso digital
Veröffentlicht: Zenodo 2024
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902297242501120
author Dae Lim, Chung
author_facet Dae Lim, Chung
contents <p>In our study, we extend the work of Vasileva et al. (1) in ”Learning Type-Aware Embeddings for Fashion Compatibility” by integrating OpenAI’s CLIP as a central component for both text and image em- beddings. This strategic replacement is premised on CLIP’s robust, pre-trained features, which encapsulate a wide range of styles and contexts, enabling more nuanced and accurate representations of fashion items. Additionally, we shift the focus towards minimizing Euclidean distances in the embedding space rather than employing the original approach of generalizing distance metrics through a fully- connected layer. This simplification aligns with the geometric intuition of embed- ding spaces, where distances more directly and transparently represent similarities and dissimilarities. Our modifications have culminated in a notable improvement in the model’s performance, enhancing the Fill In The Blank (FITB) task accuracy and Area Under Curve (AUC) metrics by 7.6% and 5% respectively. These results underscore the efficacy of our approach, marking a stride in the quest for advanced fashion recommendation systems. </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15127176
institution Zenodo
language
publishDate 2024
publisher Zenodo
record_format zenodo
spellingShingle Type Aware Embedding for Outfit Compatibility with CLIP
Dae Lim, Chung
<p>In our study, we extend the work of Vasileva et al. (1) in ”Learning Type-Aware Embeddings for Fashion Compatibility” by integrating OpenAI’s CLIP as a central component for both text and image em- beddings. This strategic replacement is premised on CLIP’s robust, pre-trained features, which encapsulate a wide range of styles and contexts, enabling more nuanced and accurate representations of fashion items. Additionally, we shift the focus towards minimizing Euclidean distances in the embedding space rather than employing the original approach of generalizing distance metrics through a fully- connected layer. This simplification aligns with the geometric intuition of embed- ding spaces, where distances more directly and transparently represent similarities and dissimilarities. Our modifications have culminated in a notable improvement in the model’s performance, enhancing the Fill In The Blank (FITB) task accuracy and Area Under Curve (AUC) metrics by 7.6% and 5% respectively. These results underscore the efficacy of our approach, marking a stride in the quest for advanced fashion recommendation systems. </p>
title Type Aware Embedding for Outfit Compatibility with CLIP
url https://doi.org/10.5281/zenodo.15127176