SafeSplit: A Novel Defense Against Client-Side Backdoor Attacks in Split Learning (Full Version)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rieger, Phillip, Pegoraro, Alessandro, Kumari, Kavita, Abera, Tigist, Knauer, Jonathan, Sadeghi, Ahmad-Reza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916627005571072
author Rieger, Phillip
Pegoraro, Alessandro
Kumari, Kavita
Abera, Tigist
Knauer, Jonathan
Sadeghi, Ahmad-Reza
author_facet Rieger, Phillip
Pegoraro, Alessandro
Kumari, Kavita
Abera, Tigist
Knauer, Jonathan
Sadeghi, Ahmad-Reza
contents Split Learning (SL) is a distributed deep learning approach enabling multiple clients and a server to collaboratively train and infer on a shared deep neural network (DNN) without requiring clients to share their private local data. The DNN is partitioned in SL, with most layers residing on the server and a few initial layers and inputs on the client side. This configuration allows resource-constrained clients to participate in training and inference. However, the distributed architecture exposes SL to backdoor attacks, where malicious clients can manipulate local datasets to alter the DNN's behavior. Existing defenses from other distributed frameworks like Federated Learning are not applicable, and there is a lack of effective backdoor defenses specifically designed for SL. We present SafeSplit, the first defense against client-side backdoor attacks in Split Learning (SL). SafeSplit enables the server to detect and filter out malicious client behavior by employing circular backward analysis after a client's training is completed, iteratively reverting to a trained checkpoint where the model under examination is found to be benign. It uses a two-fold analysis to identify client-induced changes and detect poisoned models. First, a static analysis in the frequency domain measures the differences in the layer's parameters at the server. Second, a dynamic analysis introduces a novel rotational distance metric that assesses the orientation shifts of the server's layer parameters during training. Our comprehensive evaluation across various data distributions, client counts, and attack scenarios demonstrates the high efficacy of this dual analysis in mitigating backdoor attacks while preserving model utility.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeSplit: A Novel Defense Against Client-Side Backdoor Attacks in Split Learning (Full Version)
Rieger, Phillip
Pegoraro, Alessandro
Kumari, Kavita
Abera, Tigist
Knauer, Jonathan
Sadeghi, Ahmad-Reza
Cryptography and Security
Distributed, Parallel, and Cluster Computing
Machine Learning
Split Learning (SL) is a distributed deep learning approach enabling multiple clients and a server to collaboratively train and infer on a shared deep neural network (DNN) without requiring clients to share their private local data. The DNN is partitioned in SL, with most layers residing on the server and a few initial layers and inputs on the client side. This configuration allows resource-constrained clients to participate in training and inference. However, the distributed architecture exposes SL to backdoor attacks, where malicious clients can manipulate local datasets to alter the DNN's behavior. Existing defenses from other distributed frameworks like Federated Learning are not applicable, and there is a lack of effective backdoor defenses specifically designed for SL. We present SafeSplit, the first defense against client-side backdoor attacks in Split Learning (SL). SafeSplit enables the server to detect and filter out malicious client behavior by employing circular backward analysis after a client's training is completed, iteratively reverting to a trained checkpoint where the model under examination is found to be benign. It uses a two-fold analysis to identify client-induced changes and detect poisoned models. First, a static analysis in the frequency domain measures the differences in the layer's parameters at the server. Second, a dynamic analysis introduces a novel rotational distance metric that assesses the orientation shifts of the server's layer parameters during training. Our comprehensive evaluation across various data distributions, client counts, and attack scenarios demonstrates the high efficacy of this dual analysis in mitigating backdoor attacks while preserving model utility.
title SafeSplit: A Novel Defense Against Client-Side Backdoor Attacks in Split Learning (Full Version)
topic Cryptography and Security
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2501.06650