Does CLIP perceive art the same way we do?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Asperti, Andrea, Dessì, Leonardo, Tonetti, Maria Chiara, Wu, Nico
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918174159536128
author Asperti, Andrea
Dessì, Leonardo
Tonetti, Maria Chiara
Wu, Nico
author_facet Asperti, Andrea
Dessì, Leonardo
Tonetti, Maria Chiara
Wu, Nico
contents CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate CLIP's ability to extract high-level semantic and stylistic information from paintings, including both human-created and AI-generated imagery. We evaluate its perception across multiple dimensions: content, scene understanding, artistic style, historical period, and the presence of visual deformations or artifacts. By designing targeted probing tasks and comparing CLIP's responses to human annotations and expert benchmarks, we explore its alignment with human perceptual and contextual understanding. Our findings reveal both strengths and limitations in CLIP's visual representations, particularly in relation to aesthetic cues and artistic intent. We further discuss the implications of these insights for using CLIP as a guidance mechanism during generative processes, such as style transfer or prompt-based image synthesis. Our work highlights the need for deeper interpretability in multimodal systems, especially when applied to creative domains where nuance and subjectivity play a central role.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does CLIP perceive art the same way we do?
Asperti, Andrea
Dessì, Leonardo
Tonetti, Maria Chiara
Wu, Nico
Computer Vision and Pattern Recognition
Multimedia
68T45, 68T07 (Primary) 68T50, 68U10 (Secondary)
I.2.7; I.2.10
CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate CLIP's ability to extract high-level semantic and stylistic information from paintings, including both human-created and AI-generated imagery. We evaluate its perception across multiple dimensions: content, scene understanding, artistic style, historical period, and the presence of visual deformations or artifacts. By designing targeted probing tasks and comparing CLIP's responses to human annotations and expert benchmarks, we explore its alignment with human perceptual and contextual understanding. Our findings reveal both strengths and limitations in CLIP's visual representations, particularly in relation to aesthetic cues and artistic intent. We further discuss the implications of these insights for using CLIP as a guidance mechanism during generative processes, such as style transfer or prompt-based image synthesis. Our work highlights the need for deeper interpretability in multimodal systems, especially when applied to creative domains where nuance and subjectivity play a central role.
title Does CLIP perceive art the same way we do?
topic Computer Vision and Pattern Recognition
Multimedia
68T45, 68T07 (Primary) 68T50, 68U10 (Secondary)
I.2.7; I.2.10
url https://arxiv.org/abs/2505.05229