Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Eunji, Shim, Kyuhong, Chang, Simyung, Yoon, Sungroh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910652598059008
author Kim, Eunji
Shim, Kyuhong
Chang, Simyung
Yoon, Sungroh
author_facet Kim, Eunji
Shim, Kyuhong
Chang, Simyung
Yoon, Sungroh
contents A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual elements within a sentence depending on the context, efforts to account for variation of importance in constructing text embeddings have been lacking. We propose a framework of Semantic Token Reweighting to build Interpretable text embeddings (SToRI), which incorporates controllability as well. SToRI refines the text encoding process in CLIP by differentially weighting semantic elements based on contextual importance, enabling finer control over emphasis responsive to data-driven insights and user preferences. The efficacy of SToRI is demonstrated through comprehensive experiments on few-shot image classification and image retrieval tailored to user preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
Kim, Eunji
Shim, Kyuhong
Chang, Simyung
Yoon, Sungroh
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
A text encoder within Vision-Language Models (VLMs) like CLIP plays a crucial role in translating textual input into an embedding space shared with images, thereby facilitating the interpretative analysis of vision tasks through natural language. Despite the varying significance of different textual elements within a sentence depending on the context, efforts to account for variation of importance in constructing text embeddings have been lacking. We propose a framework of Semantic Token Reweighting to build Interpretable text embeddings (SToRI), which incorporates controllability as well. SToRI refines the text encoding process in CLIP by differentially weighting semantic elements based on contextual importance, enabling finer control over emphasis responsive to data-driven insights and user preferences. The efficacy of SToRI is demonstrated through comprehensive experiments on few-shot image classification and image retrieval tailored to user preferences.
title Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.08469