GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galichin, Andrey V., Pautov, Mikhail, Zhavoronkin, Alexey, Rogov, Oleg Y., Oseledets, Ivan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910444322553856
author Galichin, Andrey V.
Pautov, Mikhail
Zhavoronkin, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
author_facet Galichin, Andrey V.
Pautov, Mikhail
Zhavoronkin, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
contents While Deep Neural Networks (DNNs) have demonstrated remarkable performance in tasks related to perception and control, there are still several unresolved concerns regarding the privacy of their training data, particularly in the context of vulnerability to Membership Inference Attacks (MIAs). In this paper, we explore a connection between the susceptibility to membership inference attacks and the vulnerability to distillation-based functionality stealing attacks. In particular, we propose {GLiRA}, a distillation-guided approach to membership inference attack on the black-box neural network. We observe that the knowledge distillation significantly improves the efficiency of likelihood ratio of membership inference attack, especially in the black-box setting, i.e., when the architecture of the target model is unknown to the attacker. We evaluate the proposed method across multiple image classification datasets and models and demonstrate that likelihood ratio attacks when guided by the knowledge distillation, outperform the current state-of-the-art membership inference attacks in the black-box setting.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07562
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
Galichin, Andrey V.
Pautov, Mikhail
Zhavoronkin, Alexey
Rogov, Oleg Y.
Oseledets, Ivan
Machine Learning
Artificial Intelligence
While Deep Neural Networks (DNNs) have demonstrated remarkable performance in tasks related to perception and control, there are still several unresolved concerns regarding the privacy of their training data, particularly in the context of vulnerability to Membership Inference Attacks (MIAs). In this paper, we explore a connection between the susceptibility to membership inference attacks and the vulnerability to distillation-based functionality stealing attacks. In particular, we propose {GLiRA}, a distillation-guided approach to membership inference attack on the black-box neural network. We observe that the knowledge distillation significantly improves the efficiency of likelihood ratio of membership inference attack, especially in the black-box setting, i.e., when the architecture of the target model is unknown to the attacker. We evaluate the proposed method across multiple image classification datasets and models and demonstrate that likelihood ratio attacks when guided by the knowledge distillation, outperform the current state-of-the-art membership inference attacks in the black-box setting.
title GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.07562