Probabilistic Stability Guarantees for Feature Attributions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jin, Helen, Xue, Anton, You, Weiqiu, Goel, Surbhi, Wong, Eric
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915433686237184
author Jin, Helen
Xue, Anton
You, Weiqiu
Goel, Surbhi
Wong, Eric
author_facet Jin, Helen
Xue, Anton
You, Weiqiu
Goel, Surbhi
Wong, Eric
contents Stability guarantees have emerged as a principled way to evaluate feature attributions, but existing certification methods rely on heavily smoothed classifiers and often produce conservative guarantees. To address these limitations, we introduce soft stability and propose a simple, model-agnostic, sample-efficient stability certification algorithm (SCA) that yields non-trivial and interpretable guarantees for any attribution method. Moreover, we show that mild smoothing achieves a more favorable trade-off between accuracy and stability, avoiding the aggressive compromises made in prior certification methods. To explain this behavior, we use Boolean function analysis to derive a novel characterization of stability under smoothing. We evaluate SCA on vision and language tasks and demonstrate the effectiveness of soft stability in measuring the robustness of explanation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probabilistic Stability Guarantees for Feature Attributions
Jin, Helen
Xue, Anton
You, Weiqiu
Goel, Surbhi
Wong, Eric
Machine Learning
Artificial Intelligence
Stability guarantees have emerged as a principled way to evaluate feature attributions, but existing certification methods rely on heavily smoothed classifiers and often produce conservative guarantees. To address these limitations, we introduce soft stability and propose a simple, model-agnostic, sample-efficient stability certification algorithm (SCA) that yields non-trivial and interpretable guarantees for any attribution method. Moreover, we show that mild smoothing achieves a more favorable trade-off between accuracy and stability, avoiding the aggressive compromises made in prior certification methods. To explain this behavior, we use Boolean function analysis to derive a novel characterization of stability under smoothing. We evaluate SCA on vision and language tasks and demonstrate the effectiveness of soft stability in measuring the robustness of explanation methods.
title Probabilistic Stability Guarantees for Feature Attributions
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.13787