Probabilistic Dataset Reconstruction from Interpretable Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferry, Julien, Aïvodji, Ulrich, Gambs, Sébastien, Huguet, Marie-José, Siala, Mohamed
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911824997253120
author Ferry, Julien
Aïvodji, Ulrich
Gambs, Sébastien
Huguet, Marie-José
Siala, Mohamed
author_facet Ferry, Julien
Aïvodji, Ulrich
Gambs, Sébastien
Huguet, Marie-José
Siala, Mohamed
contents Interpretability is often pointed out as a key requirement for trustworthy machine learning. However, learning and releasing models that are inherently interpretable leaks information regarding the underlying training data. As such disclosure may directly conflict with privacy, a precise quantification of the privacy impact of such breach is a fundamental problem. For instance, previous work have shown that the structure of a decision tree can be leveraged to build a probabilistic reconstruction of its training dataset, with the uncertainty of the reconstruction being a relevant metric for the information leak. In this paper, we propose of a novel framework generalizing these probabilistic reconstructions in the sense that it can handle other forms of interpretable models and more generic types of knowledge. In addition, we demonstrate that under realistic assumptions regarding the interpretable models' structure, the uncertainty of the reconstruction can be computed efficiently. Finally, we illustrate the applicability of our approach on both decision trees and rule lists, by comparing the theoretical information leak associated to either exact or heuristic learning algorithms. Our results suggest that optimal interpretable models are often more compact and leak less information regarding their training data than greedily-built ones, for a given accuracy level.
format Preprint
id arxiv_https___arxiv_org_abs_2308_15099
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Probabilistic Dataset Reconstruction from Interpretable Models
Ferry, Julien
Aïvodji, Ulrich
Gambs, Sébastien
Huguet, Marie-José
Siala, Mohamed
Artificial Intelligence
Information Theory
Interpretability is often pointed out as a key requirement for trustworthy machine learning. However, learning and releasing models that are inherently interpretable leaks information regarding the underlying training data. As such disclosure may directly conflict with privacy, a precise quantification of the privacy impact of such breach is a fundamental problem. For instance, previous work have shown that the structure of a decision tree can be leveraged to build a probabilistic reconstruction of its training dataset, with the uncertainty of the reconstruction being a relevant metric for the information leak. In this paper, we propose of a novel framework generalizing these probabilistic reconstructions in the sense that it can handle other forms of interpretable models and more generic types of knowledge. In addition, we demonstrate that under realistic assumptions regarding the interpretable models' structure, the uncertainty of the reconstruction can be computed efficiently. Finally, we illustrate the applicability of our approach on both decision trees and rule lists, by comparing the theoretical information leak associated to either exact or heuristic learning algorithms. Our results suggest that optimal interpretable models are often more compact and leak less information regarding their training data than greedily-built ones, for a given accuracy level.
title Probabilistic Dataset Reconstruction from Interpretable Models
topic Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2308.15099