Scaling Trends for Data Poisoning in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bowen, Dillon, Murphy, Brendan, Cai, Will, Khachaturov, David, Gleave, Adam, Pelrine, Kellin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913945240993792
author Bowen, Dillon
Murphy, Brendan
Cai, Will
Khachaturov, David
Gleave, Adam
Pelrine, Kellin
author_facet Bowen, Dillon
Murphy, Brendan
Cai, Will
Khachaturov, David
Gleave, Adam
Pelrine, Kellin
contents LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabilities in today's most capable models, this paper investigates whether these risks increase with model scaling. We evaluate three threat models -- malicious fine-tuning, imperfect data curation, and intentional data contamination -- across 24 frontier LLMs ranging from 1.5 to 72 billion parameters. Our experiments reveal that larger LLMs are significantly more susceptible to data poisoning, learning harmful behaviors from even minimal exposure to harmful data more quickly than smaller models. These findings underscore the need for leading AI companies to thoroughly red team fine-tuning APIs before public release and to develop more robust safeguards against data poisoning, particularly as models continue to scale in size and capability.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02946
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Trends for Data Poisoning in LLMs
Bowen, Dillon
Murphy, Brendan
Cai, Will
Khachaturov, David
Gleave, Adam
Pelrine, Kellin
Cryptography and Security
Artificial Intelligence
Machine Learning
LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabilities in today's most capable models, this paper investigates whether these risks increase with model scaling. We evaluate three threat models -- malicious fine-tuning, imperfect data curation, and intentional data contamination -- across 24 frontier LLMs ranging from 1.5 to 72 billion parameters. Our experiments reveal that larger LLMs are significantly more susceptible to data poisoning, learning harmful behaviors from even minimal exposure to harmful data more quickly than smaller models. These findings underscore the need for leading AI companies to thoroughly red team fine-tuning APIs before public release and to develop more robust safeguards against data poisoning, particularly as models continue to scale in size and capability.
title Scaling Trends for Data Poisoning in LLMs
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.02946