The Clever Hans Effect in Unsupervised Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kauffmann, Jacob, Dippel, Jonas, Ruff, Lukas, Samek, Wojciech, Müller, Klaus-Robert, Montavon, Grégoire
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910878387929088
author Kauffmann, Jacob
Dippel, Jonas
Ruff, Lukas
Samek, Wojciech
Müller, Klaus-Robert
Montavon, Grégoire
author_facet Kauffmann, Jacob
Dippel, Jonas
Ruff, Lukas
Samek, Wojciech
Müller, Klaus-Robert
Montavon, Grégoire
contents Unsupervised learning has become an essential building block of AI systems. The representations it produces, e.g. in foundation models, are critical to a wide variety of downstream applications. It is therefore important to carefully examine unsupervised models to ensure not only that they produce accurate predictions, but also that these predictions are not "right for the wrong reasons", the so-called Clever Hans (CH) effect. Using specially developed Explainable AI techniques, we show for the first time that CH effects are widespread in unsupervised learning. Our empirical findings are enriched by theoretical insights, which interestingly point to inductive biases in the unsupervised learning machine as a primary source of CH effects. Overall, our work sheds light on unexplored risks associated with practical applications of unsupervised learning and suggests ways to make unsupervised learning more robust.
format Preprint
id arxiv_https___arxiv_org_abs_2408_08041
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Clever Hans Effect in Unsupervised Learning
Kauffmann, Jacob
Dippel, Jonas
Ruff, Lukas
Samek, Wojciech
Müller, Klaus-Robert
Montavon, Grégoire
Machine Learning
Artificial Intelligence
Unsupervised learning has become an essential building block of AI systems. The representations it produces, e.g. in foundation models, are critical to a wide variety of downstream applications. It is therefore important to carefully examine unsupervised models to ensure not only that they produce accurate predictions, but also that these predictions are not "right for the wrong reasons", the so-called Clever Hans (CH) effect. Using specially developed Explainable AI techniques, we show for the first time that CH effects are widespread in unsupervised learning. Our empirical findings are enriched by theoretical insights, which interestingly point to inductive biases in the unsupervised learning machine as a primary source of CH effects. Overall, our work sheds light on unexplored risks associated with practical applications of unsupervised learning and suggests ways to make unsupervised learning more robust.
title The Clever Hans Effect in Unsupervised Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2408.08041