FAME: Formal Abstract Minimal Explanation for Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boumazouza, Ryma, Elsaleh, Raya, Ducoffe, Melanie, Bassan, Shahaf, Katz, Guy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918383363031040
author Boumazouza, Ryma
Elsaleh, Raya
Ducoffe, Melanie
Bassan, Shahaf
Katz, Guy
author_facet Boumazouza, Ryma
Elsaleh, Raya
Ducoffe, Melanie
Bassan, Shahaf
Katz, Guy
contents We propose FAME (Formal Abstract Minimal Explanations), a new class of abductive explanations grounded in abstract interpretation. FAME is the first method to scale to large neural networks while reducing explanation size. Our main contribution is the design of dedicated perturbation domains that eliminate the need for traversal order. FAME progressively shrinks these domains and leverages LiRPA-based bounds to discard irrelevant features, ultimately converging to a formal abstract minimal explanation. To assess explanation quality, we introduce a procedure that measures the worst-case distance between an abstract minimal explanation and a true minimal explanation. This procedure combines adversarial attacks with an optional VERIX+ refinement step. We benchmark FAME against VERIX+ and demonstrate consistent gains in both explanation size and runtime on medium- to large-scale neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10661
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FAME: Formal Abstract Minimal Explanation for Neural Networks
Boumazouza, Ryma
Elsaleh, Raya
Ducoffe, Melanie
Bassan, Shahaf
Katz, Guy
Artificial Intelligence
Machine Learning
We propose FAME (Formal Abstract Minimal Explanations), a new class of abductive explanations grounded in abstract interpretation. FAME is the first method to scale to large neural networks while reducing explanation size. Our main contribution is the design of dedicated perturbation domains that eliminate the need for traversal order. FAME progressively shrinks these domains and leverages LiRPA-based bounds to discard irrelevant features, ultimately converging to a formal abstract minimal explanation. To assess explanation quality, we introduce a procedure that measures the worst-case distance between an abstract minimal explanation and a true minimal explanation. This procedure combines adversarial attacks with an optional VERIX+ refinement step. We benchmark FAME against VERIX+ and demonstrate consistent gains in both explanation size and runtime on medium- to large-scale neural networks.
title FAME: Formal Abstract Minimal Explanation for Neural Networks
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.10661