Similarity-based Label Inference Attack against Training and Inference of Split Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Junlin, Lyu, Xinchen, Cui, Qimei, Tao, Xiaofeng
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929285336399872
author Liu, Junlin
Lyu, Xinchen
Cui, Qimei
Tao, Xiaofeng
author_facet Liu, Junlin
Lyu, Xinchen
Cui, Qimei
Tao, Xiaofeng
contents Split learning is a promising paradigm for privacy-preserving distributed learning. The learning model can be cut into multiple portions to be collaboratively trained at the participants by exchanging only the intermediate results at the cut layer. Understanding the security performance of split learning is critical for many privacy-sensitive applications. This paper shows that the exchanged intermediate results, including the smashed data (i.e., extracted features from the raw data) and gradients during training and inference of split learning, can already reveal the private labels. We mathematically analyze the potential label leakages and propose the cosine and Euclidean similarity measurements for gradients and smashed data, respectively. Then, the two similarity measurements are shown to be unified in Euclidean space. Based on the similarity metric, we design three label inference attacks to efficiently recover the private labels during both the training and inference phases. Experimental results validate that the proposed approaches can achieve close to 100% accuracy of label attacks. The proposed attack can still achieve accurate predictions against various state-of-the-art defense mechanisms, including DP-SGD, label differential privacy, gradient compression, and Marvell.
format Preprint
id arxiv_https___arxiv_org_abs_2203_05222
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Similarity-based Label Inference Attack against Training and Inference of Split Learning
Liu, Junlin
Lyu, Xinchen
Cui, Qimei
Tao, Xiaofeng
Machine Learning
Artificial Intelligence
Cryptography and Security
Split learning is a promising paradigm for privacy-preserving distributed learning. The learning model can be cut into multiple portions to be collaboratively trained at the participants by exchanging only the intermediate results at the cut layer. Understanding the security performance of split learning is critical for many privacy-sensitive applications. This paper shows that the exchanged intermediate results, including the smashed data (i.e., extracted features from the raw data) and gradients during training and inference of split learning, can already reveal the private labels. We mathematically analyze the potential label leakages and propose the cosine and Euclidean similarity measurements for gradients and smashed data, respectively. Then, the two similarity measurements are shown to be unified in Euclidean space. Based on the similarity metric, we design three label inference attacks to efficiently recover the private labels during both the training and inference phases. Experimental results validate that the proposed approaches can achieve close to 100% accuracy of label attacks. The proposed attack can still achieve accurate predictions against various state-of-the-art defense mechanisms, including DP-SGD, label differential privacy, gradient compression, and Marvell.
title Similarity-based Label Inference Attack against Training and Inference of Split Learning
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2203.05222