Causal Fuzzing for Verifying Machine Unlearning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mazhar, Anna, Galhotra, Sainyam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908549165088768
author Mazhar, Anna
Galhotra, Sainyam
author_facet Mazhar, Anna
Galhotra, Sainyam
contents As machine learning models become increasingly embedded in decision-making systems, the ability to "unlearn" targeted data or features is crucial for enhancing model adaptability, fairness, and privacy in models which involves expensive training. To effectively guide machine unlearning, a thorough testing is essential. Existing methods for verification of machine unlearning provide limited insights, often failing in scenarios where the influence is indirect. In this work, we propose CAFÉ, a new causality based framework that unifies datapoint- and feature-level unlearning for verification of black-box ML models. CAFÉ evaluates both direct and indirect effects of unlearning targets through causal dependencies, providing actionable insights with fine-grained analysis. Our evaluation across five datasets and three model architectures demonstrates that CAFÉ successfully detects residual influence missed by baselines while maintaining computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Causal Fuzzing for Verifying Machine Unlearning
Mazhar, Anna
Galhotra, Sainyam
Software Engineering
Artificial Intelligence
Machine Learning
As machine learning models become increasingly embedded in decision-making systems, the ability to "unlearn" targeted data or features is crucial for enhancing model adaptability, fairness, and privacy in models which involves expensive training. To effectively guide machine unlearning, a thorough testing is essential. Existing methods for verification of machine unlearning provide limited insights, often failing in scenarios where the influence is indirect. In this work, we propose CAFÉ, a new causality based framework that unifies datapoint- and feature-level unlearning for verification of black-box ML models. CAFÉ evaluates both direct and indirect effects of unlearning targets through causal dependencies, providing actionable insights with fine-grained analysis. Our evaluation across five datasets and three model architectures demonstrates that CAFÉ successfully detects residual influence missed by baselines while maintaining computational efficiency.
title Causal Fuzzing for Verifying Machine Unlearning
topic Software Engineering
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.16525