Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdullahu, Fitim, Grabner, Helmut
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914094094745600
author Abdullahu, Fitim
Grabner, Helmut
author_facet Abdullahu, Fitim
Grabner, Helmut
contents Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on large-scale visual and textual data has demonstrated impressive capabilities. We explore these models' potential to understand to what extent the concepts of visual interestingness are captured and examine the alignment between human assessments and GPT-4o's, a leading LMM, predictions through comparative analysis. Our studies reveal partial alignment between humans and GPT-4o. It already captures the concept as best compared to state-of-the-art methods. Hence, this allows for the effective labeling of image pairs according to their (commonly) interestingness, which are used as training data to distill the knowledge into a learning-to-rank model. The insights pave the way for a deeper understanding of human interest.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
Abdullahu, Fitim
Grabner, Helmut
Computer Vision and Pattern Recognition
Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on large-scale visual and textual data has demonstrated impressive capabilities. We explore these models' potential to understand to what extent the concepts of visual interestingness are captured and examine the alignment between human assessments and GPT-4o's, a leading LMM, predictions through comparative analysis. Our studies reveal partial alignment between humans and GPT-4o. It already captures the concept as best compared to state-of-the-art methods. Hence, this allows for the effective labeling of image pairs according to their (commonly) interestingness, which are used as training data to distill the knowledge into a learning-to-rank model. The insights pave the way for a deeper understanding of human interest.
title Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.13316