Disturbing Image Detection Using LMM-Elicited Emotion Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tzelepi, Maria, Mezaris, Vasileios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909226874437632
author Tzelepi, Maria
Mezaris, Vasileios
author_facet Tzelepi, Maria
Mezaris, Vasileios
contents In this paper we deal with the task of Disturbing Image Detection (DID), exploiting knowledge encoded in Large Multimodal Models (LMMs). Specifically, we propose to exploit LMM knowledge in a two-fold manner: first by extracting generic semantic descriptions, and second by extracting elicited emotions. Subsequently, we use the CLIP's text encoder in order to obtain the text embeddings of both the generic semantic descriptions and LMM-elicited emotions. Finally, we use the aforementioned text embeddings along with the corresponding CLIP's image embeddings for performing the DID task. The proposed method significantly improves the baseline classification accuracy, achieving state-of-the-art performance on the augmented Disturbing Image Detection dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12668
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Disturbing Image Detection Using LMM-Elicited Emotion Embeddings
Tzelepi, Maria
Mezaris, Vasileios
Computer Vision and Pattern Recognition
In this paper we deal with the task of Disturbing Image Detection (DID), exploiting knowledge encoded in Large Multimodal Models (LMMs). Specifically, we propose to exploit LMM knowledge in a two-fold manner: first by extracting generic semantic descriptions, and second by extracting elicited emotions. Subsequently, we use the CLIP's text encoder in order to obtain the text embeddings of both the generic semantic descriptions and LMM-elicited emotions. Finally, we use the aforementioned text embeddings along with the corresponding CLIP's image embeddings for performing the DID task. The proposed method significantly improves the baseline classification accuracy, achieving state-of-the-art performance on the augmented Disturbing Image Detection dataset.
title Disturbing Image Detection Using LMM-Elicited Emotion Embeddings
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.12668