Generative Classifiers Avoid Shortcut Solutions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Alexander C., Kumar, Ananya, Pathak, Deepak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918267213316096
author Li, Alexander C.
Kumar, Ananya
Pathak, Deepak
author_facet Li, Alexander C.
Kumar, Ananya
Pathak, Deepak
contents Discriminative approaches to classification often learn shortcuts that hold in-distribution but fail even under minor distribution shift. This failure mode stems from an overreliance on features that are spuriously correlated with the label. We show that generative classifiers, which use class-conditional generative models, can avoid this issue by modeling all features, both core and spurious, instead of mainly spurious ones. These generative classifiers are simple to train, avoiding the need for specialized augmentations, strong regularization, extra hyperparameters, or knowledge of the specific spurious correlations to avoid. We find that diffusion-based and autoregressive generative classifiers achieve state-of-the-art performance on five standard image and text distribution shift benchmarks and reduce the impact of spurious correlations in realistic applications, such as medical or satellite datasets. Finally, we carefully analyze a Gaussian toy setting to understand the inductive biases of generative classifiers, as well as the data properties that determine when generative classifiers outperform discriminative ones.
format Preprint
id arxiv_https___arxiv_org_abs_2512_25034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Classifiers Avoid Shortcut Solutions
Li, Alexander C.
Kumar, Ananya
Pathak, Deepak
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
Discriminative approaches to classification often learn shortcuts that hold in-distribution but fail even under minor distribution shift. This failure mode stems from an overreliance on features that are spuriously correlated with the label. We show that generative classifiers, which use class-conditional generative models, can avoid this issue by modeling all features, both core and spurious, instead of mainly spurious ones. These generative classifiers are simple to train, avoiding the need for specialized augmentations, strong regularization, extra hyperparameters, or knowledge of the specific spurious correlations to avoid. We find that diffusion-based and autoregressive generative classifiers achieve state-of-the-art performance on five standard image and text distribution shift benchmarks and reduce the impact of spurious correlations in realistic applications, such as medical or satellite datasets. Finally, we carefully analyze a Gaussian toy setting to understand the inductive biases of generative classifiers, as well as the data properties that determine when generative classifiers outperform discriminative ones.
title Generative Classifiers Avoid Shortcut Solutions
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Neural and Evolutionary Computing
url https://arxiv.org/abs/2512.25034