Saved in:
Bibliographic Details
Main Authors: Méndez, David, Bontempo, Gianpaolo, Ficarra, Elisa, Confalonieri, Roberto, Díaz-Rodríguez, Natalia
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.11060
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916739753705472
author Méndez, David
Bontempo, Gianpaolo
Ficarra, Elisa
Confalonieri, Roberto
Díaz-Rodríguez, Natalia
author_facet Méndez, David
Bontempo, Gianpaolo
Ficarra, Elisa
Confalonieri, Roberto
Díaz-Rodríguez, Natalia
contents Deep vision models often rely on biases learned from spurious correlations in datasets. To identify these biases, methods that interpret high-level, human-understandable concepts are more effective than those relying primarily on low-level features like heatmaps. A major challenge for these concept-based methods is the lack of image annotations indicating potentially bias-inducing concepts, since creating such annotations requires detailed labeling for each dataset and concept, which is highly labor-intensive. We present CUBIC (Concept embeddings for Unsupervised Bias IdentifiCation), a novel method that automatically discovers interpretable concepts that may bias classifier behavior. Unlike existing approaches, CUBIC does not rely on predefined bias candidates or examples of model failures tied to specific biases, as such information is not always available. Instead, it leverages image-text latent space and linear classifier probes to examine how the latent representation of a superclass label$\unicode{x2014}$shared by all instances in the dataset$\unicode{x2014}$is influenced by the presence of a given concept. By measuring these shifts against the normal vector to the classifier's decision boundary, CUBIC identifies concepts that significantly influence model predictions. Our experiments demonstrate that CUBIC effectively uncovers previously unknown biases using Vision-Language Models (VLMs) without requiring the samples in the dataset where the classifier underperforms or prior knowledge of potential biases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
Méndez, David
Bontempo, Gianpaolo
Ficarra, Elisa
Confalonieri, Roberto
Díaz-Rodríguez, Natalia
Computer Vision and Pattern Recognition
Artificial Intelligence
68T10
I.2.4; I.5.2
Deep vision models often rely on biases learned from spurious correlations in datasets. To identify these biases, methods that interpret high-level, human-understandable concepts are more effective than those relying primarily on low-level features like heatmaps. A major challenge for these concept-based methods is the lack of image annotations indicating potentially bias-inducing concepts, since creating such annotations requires detailed labeling for each dataset and concept, which is highly labor-intensive. We present CUBIC (Concept embeddings for Unsupervised Bias IdentifiCation), a novel method that automatically discovers interpretable concepts that may bias classifier behavior. Unlike existing approaches, CUBIC does not rely on predefined bias candidates or examples of model failures tied to specific biases, as such information is not always available. Instead, it leverages image-text latent space and linear classifier probes to examine how the latent representation of a superclass label$\unicode{x2014}$shared by all instances in the dataset$\unicode{x2014}$is influenced by the presence of a given concept. By measuring these shifts against the normal vector to the classifier's decision boundary, CUBIC identifies concepts that significantly influence model predictions. Our experiments demonstrate that CUBIC effectively uncovers previously unknown biases using Vision-Language Models (VLMs) without requiring the samples in the dataset where the classifier underperforms or prior knowledge of potential biases.
title CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
68T10
I.2.4; I.5.2
url https://arxiv.org/abs/2505.11060