If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Hernandez, Adriano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913421840089088
author Hernandez, Adriano
author_facet Hernandez, Adriano
contents Large language models (LLMs) sometimes exhibit dangerous unintended behaviors. Finding and fixing these is challenging because the attack surface is massive -- it is not tractable to exhaustively search for all possible inputs that may elicit such behavior. One specific and particularly challenging case is that if data-poisoning-injected trojans, since there is no way to know what they are to search for them. To our knowledge, there is no generally applicable method to unlearn unknown trojans injected during pre-training. This work seeks to provide a general purpose recipe (filters) and a specific implementation (LoRA) filters that work in practice on small to medium sized models. The focus is primarily empirical, though some perplexing behavior opens the door to the fundamental question of how LLMs store and process information. Not unexpectedly, we find that our filters work best on the residual stream and the latest layers.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
Hernandez, Adriano
Machine Learning
Computation and Language
Cryptography and Security
Large language models (LLMs) sometimes exhibit dangerous unintended behaviors. Finding and fixing these is challenging because the attack surface is massive -- it is not tractable to exhaustively search for all possible inputs that may elicit such behavior. One specific and particularly challenging case is that if data-poisoning-injected trojans, since there is no way to know what they are to search for them. To our knowledge, there is no generally applicable method to unlearn unknown trojans injected during pre-training. This work seeks to provide a general purpose recipe (filters) and a specific implementation (LoRA) filters that work in practice on small to medium sized models. The focus is primarily empirical, though some perplexing behavior opens the door to the fundamental question of how LLMs store and process information. Not unexpectedly, we find that our filters work best on the residual stream and the latest layers.
title If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2407.06411