Counterfactual Stress Testing for Image Classification Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stammel, Moritz, Ribeiro, Fabio De Sousa, Mehta, Raghav, Roschewitz, Mélanie, Glocker, Ben
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918494760599552
author Stammel, Moritz
Ribeiro, Fabio De Sousa
Mehta, Raghav
Roschewitz, Mélanie
Glocker, Ben
author_facet Stammel, Moritz
Ribeiro, Fabio De Sousa
Mehta, Raghav
Roschewitz, Mélanie
Glocker, Ben
contents Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardware, or acquisition protocols. A central challenge is underspecification, where models with similar validation performance exhibit divergent real-world failure modes. Although stress testing has emerged as a tool to assess this, current methods typically rely on simple, uninformed perturbations (e.g., brightness or contrast changes), which fail to capture clinically realistic variation and can overestimate robustness. In this work, we introduce a counterfactual stress testing framework based on causal generative models that create realistic "what if" images by intervening on attributes such as scanner type and patient sex while preserving anatomical identity, enabling controlled and semantically meaningful evaluation under targeted distribution shifts. Across two imaging modalities (chest X-ray and mammography), three model architectures, and multiple shift scenarios, we show that counterfactual stress tests provide a substantially more accurate proxy for real out-of-distribution performance than classical perturbations, capturing the direction and relative magnitude of performance changes as well as model ranking. These results suggest that causal generative models can serve as practical simulators for robustness assessment, offering a more reliable basis for evaluating medical AI systems prior to deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10894
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Counterfactual Stress Testing for Image Classification Models
Stammel, Moritz
Ribeiro, Fabio De Sousa
Mehta, Raghav
Roschewitz, Mélanie
Glocker, Ben
Computer Vision and Pattern Recognition
Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardware, or acquisition protocols. A central challenge is underspecification, where models with similar validation performance exhibit divergent real-world failure modes. Although stress testing has emerged as a tool to assess this, current methods typically rely on simple, uninformed perturbations (e.g., brightness or contrast changes), which fail to capture clinically realistic variation and can overestimate robustness. In this work, we introduce a counterfactual stress testing framework based on causal generative models that create realistic "what if" images by intervening on attributes such as scanner type and patient sex while preserving anatomical identity, enabling controlled and semantically meaningful evaluation under targeted distribution shifts. Across two imaging modalities (chest X-ray and mammography), three model architectures, and multiple shift scenarios, we show that counterfactual stress tests provide a substantially more accurate proxy for real out-of-distribution performance than classical perturbations, capturing the direction and relative magnitude of performance changes as well as model ranking. These results suggest that causal generative models can serve as practical simulators for robustness assessment, offering a more reliable basis for evaluating medical AI systems prior to deployment.
title Counterfactual Stress Testing for Image Classification Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10894