Unveiling the Unseen: Exploring Whitebox Membership Inference through the Lens of Explainability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenxi, Kumar, Abhinav, Guo, Zhen, Hou, Jie, Tourani, Reza
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909236347273216
author Li, Chenxi
Kumar, Abhinav
Guo, Zhen
Hou, Jie
Tourani, Reza
author_facet Li, Chenxi
Kumar, Abhinav
Guo, Zhen
Hou, Jie
Tourani, Reza
contents The increasing prominence of deep learning applications and reliance on personalized data underscore the urgent need to address privacy vulnerabilities, particularly Membership Inference Attacks (MIAs). Despite numerous MIA studies, significant knowledge gaps persist, particularly regarding the impact of hidden features (in isolation) on attack efficacy and insufficient justification for the root causes of attacks based on raw data features. In this paper, we aim to address these knowledge gaps by first exploring statistical approaches to identify the most informative neurons and quantifying the significance of the hidden activations from the selected neurons on attack accuracy, in isolation and combination. Additionally, we propose an attack-driven explainable framework by integrating the target and attack models to identify the most influential features of raw data that lead to successful membership inference attacks. Our proposed MIA shows an improvement of up to 26% on state-of-the-art MIA.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01306
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Unseen: Exploring Whitebox Membership Inference through the Lens of Explainability
Li, Chenxi
Kumar, Abhinav
Guo, Zhen
Hou, Jie
Tourani, Reza
Machine Learning
Cryptography and Security
The increasing prominence of deep learning applications and reliance on personalized data underscore the urgent need to address privacy vulnerabilities, particularly Membership Inference Attacks (MIAs). Despite numerous MIA studies, significant knowledge gaps persist, particularly regarding the impact of hidden features (in isolation) on attack efficacy and insufficient justification for the root causes of attacks based on raw data features. In this paper, we aim to address these knowledge gaps by first exploring statistical approaches to identify the most informative neurons and quantifying the significance of the hidden activations from the selected neurons on attack accuracy, in isolation and combination. Additionally, we propose an attack-driven explainable framework by integrating the target and attack models to identify the most influential features of raw data that lead to successful membership inference attacks. Our proposed MIA shows an improvement of up to 26% on state-of-the-art MIA.
title Unveiling the Unseen: Exploring Whitebox Membership Inference through the Lens of Explainability
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2407.01306