Trained Random Forests Completely Reveal your Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ferry, Julien, Fukasawa, Ricardo, Pascal, Timothée, Vidal, Thibaut
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909287240957952
author Ferry, Julien
Fukasawa, Ricardo
Pascal, Timothée
Vidal, Thibaut
author_facet Ferry, Julien
Fukasawa, Ricardo
Pascal, Timothée
Vidal, Thibaut
contents We introduce an optimization-based reconstruction attack capable of completely or near-completely reconstructing a dataset utilized for training a random forest. Notably, our approach relies solely on information readily available in commonly used libraries such as scikit-learn. To achieve this, we formulate the reconstruction problem as a combinatorial problem under a maximum likelihood objective. We demonstrate that this problem is NP-hard, though solvable at scale using constraint programming -- an approach rooted in constraint propagation and solution-domain reduction. Through an extensive computational investigation, we demonstrate that random forests trained without bootstrap aggregation but with feature randomization are susceptible to a complete reconstruction. This holds true even with a small number of trees. Even with bootstrap aggregation, the majority of the data can also be reconstructed. These findings underscore a critical vulnerability inherent in widely adopted ensemble methods, warranting attention and mitigation. Although the potential for such reconstruction attacks has been discussed in privacy research, our study provides clear empirical evidence of their practicability.
format Preprint
id arxiv_https___arxiv_org_abs_2402_19232
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Trained Random Forests Completely Reveal your Dataset
Ferry, Julien
Fukasawa, Ricardo
Pascal, Timothée
Vidal, Thibaut
Machine Learning
Cryptography and Security
We introduce an optimization-based reconstruction attack capable of completely or near-completely reconstructing a dataset utilized for training a random forest. Notably, our approach relies solely on information readily available in commonly used libraries such as scikit-learn. To achieve this, we formulate the reconstruction problem as a combinatorial problem under a maximum likelihood objective. We demonstrate that this problem is NP-hard, though solvable at scale using constraint programming -- an approach rooted in constraint propagation and solution-domain reduction. Through an extensive computational investigation, we demonstrate that random forests trained without bootstrap aggregation but with feature randomization are susceptible to a complete reconstruction. This holds true even with a small number of trees. Even with bootstrap aggregation, the majority of the data can also be reconstructed. These findings underscore a critical vulnerability inherent in widely adopted ensemble methods, warranting attention and mitigation. Although the potential for such reconstruction attacks has been discussed in privacy research, our study provides clear empirical evidence of their practicability.
title Trained Random Forests Completely Reveal your Dataset
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2402.19232