Poison to Detect: Detection of Targeted Overfitting in Federated Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mestari, Soumia Zohra El, Zuziak, Maciej Krzysztof, Lenzini, Gabriele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911155556974592
author Mestari, Soumia Zohra El
Zuziak, Maciej Krzysztof
Lenzini, Gabriele
author_facet Mestari, Soumia Zohra El
Zuziak, Maciej Krzysztof
Lenzini, Gabriele
contents Federated Learning (FL) enables collaborative model training across decentralised clients while keeping local data private, making it a widely adopted privacy-enhancing technology (PET). Despite its privacy benefits, FL remains vulnerable to privacy attacks, including those targeting specific clients. In this paper, we study an underexplored threat where a dishonest orchestrator intentionally manipulates the aggregation process to induce targeted overfitting in the local models of specific clients. Whereas many studies in this area predominantly focus on reducing the amount of information leakage during training, we focus on enabling an early client-side detection of targeted overfitting, thereby allowing clients to disengage before significant harm occurs. In line with this, we propose three detection techniques - (a) label flipping, (b) backdoor trigger injection, and (c) model fingerprinting - that enable clients to verify the integrity of the global aggregation. We evaluated our methods on multiple datasets under different attack scenarios. Our results show that the three methods reliably detect targeted overfitting induced by the orchestrator, but they differ in terms of computational complexity, detection latency, and false-positive rates.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11974
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Poison to Detect: Detection of Targeted Overfitting in Federated Learning
Mestari, Soumia Zohra El
Zuziak, Maciej Krzysztof
Lenzini, Gabriele
Cryptography and Security
Artificial Intelligence
Federated Learning (FL) enables collaborative model training across decentralised clients while keeping local data private, making it a widely adopted privacy-enhancing technology (PET). Despite its privacy benefits, FL remains vulnerable to privacy attacks, including those targeting specific clients. In this paper, we study an underexplored threat where a dishonest orchestrator intentionally manipulates the aggregation process to induce targeted overfitting in the local models of specific clients. Whereas many studies in this area predominantly focus on reducing the amount of information leakage during training, we focus on enabling an early client-side detection of targeted overfitting, thereby allowing clients to disengage before significant harm occurs. In line with this, we propose three detection techniques - (a) label flipping, (b) backdoor trigger injection, and (c) model fingerprinting - that enable clients to verify the integrity of the global aggregation. We evaluated our methods on multiple datasets under different attack scenarios. Our results show that the three methods reliably detect targeted overfitting induced by the orchestrator, but they differ in terms of computational complexity, detection latency, and false-positive rates.
title Poison to Detect: Detection of Targeted Overfitting in Federated Learning
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2509.11974