EmoGist: Efficient In-Context Learning for Visual Emotion Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Seoh, Ronald, Goldwasser, Dan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912595062030336
author Seoh, Ronald
Goldwasser, Dan
author_facet Seoh, Ronald
Goldwasser, Dan
contents In this paper, we introduce EmoGist, a training-free, in-context learning method for performing visual emotion classification with LVLMs. The key intuition of our approach is that context-dependent definition of emotion labels could allow more accurate predictions of emotions, as the ways in which emotions manifest within images are highly context dependent and nuanced. EmoGist pre-generates multiple descriptions of emotion labels, by analyzing the clusters of example images belonging to each label. At test time, we retrieve a version of description based on the cosine similarity of test image to cluster centroids, and feed it together with the test image to a fast LVLM for classification. Through our experiments, we show that EmoGist allows up to 12 points improvement in micro F1 scores with the multi-label Memotion dataset, and up to 8 points in macro F1 in the multi-class FI dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14660
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
Seoh, Ronald
Goldwasser, Dan
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
In this paper, we introduce EmoGist, a training-free, in-context learning method for performing visual emotion classification with LVLMs. The key intuition of our approach is that context-dependent definition of emotion labels could allow more accurate predictions of emotions, as the ways in which emotions manifest within images are highly context dependent and nuanced. EmoGist pre-generates multiple descriptions of emotion labels, by analyzing the clusters of example images belonging to each label. At test time, we retrieve a version of description based on the cosine similarity of test image to cluster centroids, and feed it together with the test image to a fast LVLM for classification. Through our experiments, we show that EmoGist allows up to 12 points improvement in micro F1 scores with the multi-label Memotion dataset, and up to 8 points in macro F1 in the multi-class FI dataset.
title EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14660