Bridging Explainability and Embeddings: BEE Aware of Spuriousness

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Păduraru, Cristian Daniel, Bărbălau, Antonio, Filipescu, Radu, Nicolicioiu, Andrei Liviu, Burceanu, Elena
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918331514093568
author Păduraru, Cristian Daniel
Bărbălau, Antonio
Filipescu, Radu
Nicolicioiu, Andrei Liviu
Burceanu, Elena
author_facet Păduraru, Cristian Daniel
Bărbălau, Antonio
Filipescu, Radu
Nicolicioiu, Andrei Liviu
Burceanu, Elena
contents Current methods for detecting spurious correlations rely on analyzing dataset statistics or error patterns, leaving many harmful shortcuts invisible when counterexamples are absent. We introduce BEE (Bridging Explainability and Embeddings), a framework that shifts the focus from model predictions to the weight space, and to the embedding geometry underlying decisions. By analyzing how fine-tuning perturbs pretrained representations, BEE uncovers spurious correlations that remain hidden from conventional evaluation pipelines. We use linear probing as a transparent diagnostic lens, revealing spurious features that not only persist after full fine-tuning but also transfer across diverse state-of-the-art models. Our experiments cover numerous datasets and domains: vision (Waterbirds, CelebA, ImageNet-1k), language (CivilComments, MIMIC-CXR medical notes), and multiple embedding families (CLIP, CLIP-DataComp.XL, mGTE, BLIP2, SigLIP2). BEE consistently exposes spurious correlations: from concepts that slash the ImageNet accuracy by up to 95%, to clinical shortcuts in MIMIC-CXR notes that induce dangerous false negatives. Together, these results position BEE as a general and principled tool for diagnosing spurious correlations in weight space, enabling principled dataset auditing and more trustworthy foundation models. The source code is publicly available at https://github.com/bit-ml/bee.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18970
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging Explainability and Embeddings: BEE Aware of Spuriousness
Păduraru, Cristian Daniel
Bărbălau, Antonio
Filipescu, Radu
Nicolicioiu, Andrei Liviu
Burceanu, Elena
Artificial Intelligence
Machine Learning
Current methods for detecting spurious correlations rely on analyzing dataset statistics or error patterns, leaving many harmful shortcuts invisible when counterexamples are absent. We introduce BEE (Bridging Explainability and Embeddings), a framework that shifts the focus from model predictions to the weight space, and to the embedding geometry underlying decisions. By analyzing how fine-tuning perturbs pretrained representations, BEE uncovers spurious correlations that remain hidden from conventional evaluation pipelines. We use linear probing as a transparent diagnostic lens, revealing spurious features that not only persist after full fine-tuning but also transfer across diverse state-of-the-art models. Our experiments cover numerous datasets and domains: vision (Waterbirds, CelebA, ImageNet-1k), language (CivilComments, MIMIC-CXR medical notes), and multiple embedding families (CLIP, CLIP-DataComp.XL, mGTE, BLIP2, SigLIP2). BEE consistently exposes spurious correlations: from concepts that slash the ImageNet accuracy by up to 95%, to clinical shortcuts in MIMIC-CXR notes that induce dangerous false negatives. Together, these results position BEE as a general and principled tool for diagnosing spurious correlations in weight space, enabling principled dataset auditing and more trustworthy foundation models. The source code is publicly available at https://github.com/bit-ml/bee.
title Bridging Explainability and Embeddings: BEE Aware of Spuriousness
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.18970