Exploiting Supervised Poison Vulnerability to Strengthen Self-Supervised Defense

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Styborski, Jeremy, Lyu, Mingzhi, Huang, Yi, Kong, Adams
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913498896793600
author Styborski, Jeremy
Lyu, Mingzhi
Huang, Yi
Kong, Adams
author_facet Styborski, Jeremy
Lyu, Mingzhi
Huang, Yi
Kong, Adams
contents Availability poisons exploit supervised learning (SL) algorithms by introducing class-related shortcut features in images such that models trained on poisoned data are useless for real-world datasets. Self-supervised learning (SSL), which utilizes augmentations to learn instance discrimination, is regarded as a strong defense against poisoned data. However, by extending the study of SSL across multiple poisons on the CIFAR-10 and ImageNet-100 datasets, we demonstrate that it often performs poorly, far below that of training on clean data. Leveraging the vulnerability of SL to poison attacks, we introduce adversarial training (AT) on SL to obfuscate poison features and guide robust feature learning for SSL. Our proposed defense, designated VESPR (Vulnerability Exploitation of Supervised Poisoning for Robust SSL), surpasses the performance of six previous defenses across seven popular availability poisons. VESPR displays superior performance over all previous defenses, boosting the minimum and average ImageNet-100 test accuracies of poisoned models by 16% and 9%, respectively. Through analysis and ablation studies, we elucidate the mechanisms by which VESPR learns robust class features.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08509
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting Supervised Poison Vulnerability to Strengthen Self-Supervised Defense
Styborski, Jeremy
Lyu, Mingzhi
Huang, Yi
Kong, Adams
Computer Vision and Pattern Recognition
Availability poisons exploit supervised learning (SL) algorithms by introducing class-related shortcut features in images such that models trained on poisoned data are useless for real-world datasets. Self-supervised learning (SSL), which utilizes augmentations to learn instance discrimination, is regarded as a strong defense against poisoned data. However, by extending the study of SSL across multiple poisons on the CIFAR-10 and ImageNet-100 datasets, we demonstrate that it often performs poorly, far below that of training on clean data. Leveraging the vulnerability of SL to poison attacks, we introduce adversarial training (AT) on SL to obfuscate poison features and guide robust feature learning for SSL. Our proposed defense, designated VESPR (Vulnerability Exploitation of Supervised Poisoning for Robust SSL), surpasses the performance of six previous defenses across seven popular availability poisons. VESPR displays superior performance over all previous defenses, boosting the minimum and average ImageNet-100 test accuracies of poisoned models by 16% and 9%, respectively. Through analysis and ablation studies, we elucidate the mechanisms by which VESPR learns robust class features.
title Exploiting Supervised Poison Vulnerability to Strengthen Self-Supervised Defense
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.08509