SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xue, Eric, Zhang, Ruiyi, Xie, Pengtao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908748133433344
author Xue, Eric
Zhang, Ruiyi
Xie, Pengtao
author_facet Xue, Eric
Zhang, Ruiyi
Xie, Pengtao
contents Modern language models remain vulnerable to backdoor attacks via poisoned data, where training inputs containing a trigger are paired with a target output, causing the model to reproduce that behavior whenever the trigger appears at inference time. Recent work has emphasized stealthy attacks that stress-test data-curation defenses using stylized artifacts or token-level perturbations as triggers, but this focus leaves a more practically relevant threat model underexplored: backdoors tied to naturally occurring semantic concepts. We introduce SteganoBackdoor, an optimization-based framework that constructs SteganoPoisons, steganographic poisoned training examples in which a backdoor payload is distributed across a fluent sentence while exhibiting no representational overlap with the inference-time semantic trigger. Across diverse model architectures, SteganoBackdoor achieves high attack success under constrained poisoning budgets and remains effective under conservative data-level filtering, highlighting a blind spot in existing data-curation defenses.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
Xue, Eric
Zhang, Ruiyi
Xie, Pengtao
Cryptography and Security
Computation and Language
Machine Learning
Modern language models remain vulnerable to backdoor attacks via poisoned data, where training inputs containing a trigger are paired with a target output, causing the model to reproduce that behavior whenever the trigger appears at inference time. Recent work has emphasized stealthy attacks that stress-test data-curation defenses using stylized artifacts or token-level perturbations as triggers, but this focus leaves a more practically relevant threat model underexplored: backdoors tied to naturally occurring semantic concepts. We introduce SteganoBackdoor, an optimization-based framework that constructs SteganoPoisons, steganographic poisoned training examples in which a backdoor payload is distributed across a fluent sentence while exhibiting no representational overlap with the inference-time semantic trigger. Across diverse model architectures, SteganoBackdoor achieves high attack success under constrained poisoning budgets and remains effective under conservative data-level filtering, highlighting a blind spot in existing data-curation defenses.
title SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
topic Cryptography and Security
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.14301