Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: De Muri, Giovanni, Vero, Mark, Staab, Robin, Vechev, Martin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914105258934272
author De Muri, Giovanni
Vero, Mark
Staab, Robin
Vechev, Martin
author_facet De Muri, Giovanni
Vero, Mark
Staab, Robin
Vechev, Martin
contents LLMs are often used by downstream users as teacher models for knowledge distillation, compressing their capabilities into memory-efficient models. However, as these teacher models may stem from untrusted parties, distillation can raise unexpected security risks. In this paper, we investigate the security implications of knowledge distillation from backdoored teacher models. First, we show that prior backdoors mostly do not transfer onto student models. Our key insight is that this is because existing LLM backdooring methods choose trigger tokens that rarely occur in usual contexts. We argue that this underestimates the security risks of knowledge distillation and introduce a new backdooring technique, T-MTB, that enables the construction and study of transferable backdoors. T-MTB carefully constructs a composite backdoor trigger, made up of several specific tokens that often occur individually in anticipated distillation datasets. As such, the poisoned teacher remains stealthy, while during distillation the individual presence of these tokens provides enough signal for the backdoor to transfer onto the student. Using T-MTB, we demonstrate and extensively study the security risks of transferable backdoors across two attack scenarios, jailbreaking and content modulation, and across four model families of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
De Muri, Giovanni
Vero, Mark
Staab, Robin
Vechev, Martin
Machine Learning
Artificial Intelligence
Cryptography and Security
LLMs are often used by downstream users as teacher models for knowledge distillation, compressing their capabilities into memory-efficient models. However, as these teacher models may stem from untrusted parties, distillation can raise unexpected security risks. In this paper, we investigate the security implications of knowledge distillation from backdoored teacher models. First, we show that prior backdoors mostly do not transfer onto student models. Our key insight is that this is because existing LLM backdooring methods choose trigger tokens that rarely occur in usual contexts. We argue that this underestimates the security risks of knowledge distillation and introduce a new backdooring technique, T-MTB, that enables the construction and study of transferable backdoors. T-MTB carefully constructs a composite backdoor trigger, made up of several specific tokens that often occur individually in anticipated distillation datasets. As such, the poisoned teacher remains stealthy, while during distillation the individual presence of these tokens provides enough signal for the backdoor to transfer onto the student. Using T-MTB, we demonstrate and extensively study the security risks of transferable backdoors across two attack scenarios, jailbreaking and content modulation, and across four model families of LLMs.
title Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2510.18541