Unlearning-based Neural Interpretations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Choi, Ching Lam, Duplessis, Alexandre, Belongie, Serge
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910821274091520
author Choi, Ching Lam
Duplessis, Alexandre
Belongie, Serge
author_facet Choi, Ching Lam
Duplessis, Alexandre
Belongie, Serge
contents Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static functions--constant mapping, averaging or blurring--inject harmful colour, texture or frequency assumptions that deviate from model behaviour. This leads to accumulation of irregular gradients, resulting in attribution maps that are biased, fragile and manipulable. Departing from the static approach, we propose UNI to compute an (un)learnable, debiased and adaptive baseline by perturbing the input towards an unlearning direction of steepest ascent. Our method discovers reliable baselines and succeeds in erasing salient features, which in turn locally smooths the high-curvature decision boundaries. Our analyses point to unlearning as a promising avenue for generating faithful, efficient and robust interpretations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08069
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unlearning-based Neural Interpretations
Choi, Ching Lam
Duplessis, Alexandre
Belongie, Serge
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static functions--constant mapping, averaging or blurring--inject harmful colour, texture or frequency assumptions that deviate from model behaviour. This leads to accumulation of irregular gradients, resulting in attribution maps that are biased, fragile and manipulable. Departing from the static approach, we propose UNI to compute an (un)learnable, debiased and adaptive baseline by perturbing the input towards an unlearning direction of steepest ascent. Our method discovers reliable baselines and succeeds in erasing salient features, which in turn locally smooths the high-curvature decision boundaries. Our analyses point to unlearning as a promising avenue for generating faithful, efficient and robust interpretations.
title Unlearning-based Neural Interpretations
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.08069