Poison in the Well: Feature Embedding Disruption in Backdoor Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Zhou, Chen, Jiahao, Zhou, Chunyi, Pu, Yuwen, Li, Qingming, Ji, Shouling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916759373611008
author Feng, Zhou
Chen, Jiahao
Zhou, Chunyi
Pu, Yuwen
Li, Qingming
Ji, Shouling
author_facet Feng, Zhou
Chen, Jiahao
Zhou, Chunyi
Pu, Yuwen
Li, Qingming
Ji, Shouling
contents Backdoor attacks embed malicious triggers into training data, enabling attackers to manipulate neural network behavior during inference while maintaining high accuracy on benign inputs. However, existing backdoor attacks face limitations manifesting in excessive reliance on training data, poor stealth, and instability, which hinder their effectiveness in real-world applications. Therefore, this paper introduces ShadowPrint, a versatile backdoor attack that targets feature embeddings within neural networks to achieve high ASRs and stealthiness. Unlike traditional approaches, ShadowPrint reduces reliance on training data access and operates effectively with exceedingly low poison rates (as low as 0.01%). It leverages a clustering-based optimization strategy to align feature embeddings, ensuring robust performance across diverse scenarios while maintaining stability and stealth. Extensive evaluations demonstrate that ShadowPrint achieves superior ASR (up to 100%), steady CA (with decay no more than 1% in most cases), and low DDR (averaging below 5%) across both clean-label and dirty-label settings, and with poison rates ranging from as low as 0.01% to 0.05%, setting a new standard for backdoor attack capabilities and emphasizing the need for advanced defense strategies focused on feature space manipulations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
Feng, Zhou
Chen, Jiahao
Zhou, Chunyi
Pu, Yuwen
Li, Qingming
Ji, Shouling
Cryptography and Security
Machine Learning
I.2.6; I.5.1; D.4.6
Backdoor attacks embed malicious triggers into training data, enabling attackers to manipulate neural network behavior during inference while maintaining high accuracy on benign inputs. However, existing backdoor attacks face limitations manifesting in excessive reliance on training data, poor stealth, and instability, which hinder their effectiveness in real-world applications. Therefore, this paper introduces ShadowPrint, a versatile backdoor attack that targets feature embeddings within neural networks to achieve high ASRs and stealthiness. Unlike traditional approaches, ShadowPrint reduces reliance on training data access and operates effectively with exceedingly low poison rates (as low as 0.01%). It leverages a clustering-based optimization strategy to align feature embeddings, ensuring robust performance across diverse scenarios while maintaining stability and stealth. Extensive evaluations demonstrate that ShadowPrint achieves superior ASR (up to 100%), steady CA (with decay no more than 1% in most cases), and low DDR (averaging below 5%) across both clean-label and dirty-label settings, and with poison rates ranging from as low as 0.01% to 0.05%, setting a new standard for backdoor attack capabilities and emphasizing the need for advanced defense strategies focused on feature space manipulations.
title Poison in the Well: Feature Embedding Disruption in Backdoor Attacks
topic Cryptography and Security
Machine Learning
I.2.6; I.5.1; D.4.6
url https://arxiv.org/abs/2505.19821