Trojan Cleansing with Neural Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Xihe, Fields, Greg, Jandali, Yaman, Javidi, Tara, Koushanfar, Farinaz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912762850967552
author Gu, Xihe
Fields, Greg
Jandali, Yaman
Javidi, Tara
Koushanfar, Farinaz
author_facet Gu, Xihe
Fields, Greg
Jandali, Yaman
Javidi, Tara
Koushanfar, Farinaz
contents Trojan attacks are sophisticated training-time attacks on neural networks that embed backdoor triggers which force the network to produce a specific output on any input which includes the trigger. With the increasing relevance of deep networks which are too large to train with personal resources and which are trained on data too large to thoroughly audit, these training-time attacks pose a significant risk. In this work, we connect trojan attacks to Neural Collapse, a phenomenon wherein the final feature representations of over-parameterized neural networks converge to a simple geometric structure. We provide experimental evidence that trojan attacks disrupt this convergence for a variety of datasets and architectures. We then use this disruption to design a lightweight, broadly generalizable mechanism for cleansing trojan attacks from a wide variety of different network architectures and experimentally demonstrate its efficacy.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12914
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Trojan Cleansing with Neural Collapse
Gu, Xihe
Fields, Greg
Jandali, Yaman
Javidi, Tara
Koushanfar, Farinaz
Machine Learning
Cryptography and Security
Trojan attacks are sophisticated training-time attacks on neural networks that embed backdoor triggers which force the network to produce a specific output on any input which includes the trigger. With the increasing relevance of deep networks which are too large to train with personal resources and which are trained on data too large to thoroughly audit, these training-time attacks pose a significant risk. In this work, we connect trojan attacks to Neural Collapse, a phenomenon wherein the final feature representations of over-parameterized neural networks converge to a simple geometric structure. We provide experimental evidence that trojan attacks disrupt this convergence for a variety of datasets and architectures. We then use this disruption to design a lightweight, broadly generalizable mechanism for cleansing trojan attacks from a wide variety of different network architectures and experimentally demonstrate its efficacy.
title Trojan Cleansing with Neural Collapse
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2411.12914