Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mamun, Md Abdullah Al, Alouani, Ihsen, Abu-Ghazaleh, Nael
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915467745034240
author Mamun, Md Abdullah Al
Alouani, Ihsen
Abu-Ghazaleh, Nael
author_facet Mamun, Md Abdullah Al
Alouani, Ihsen
Abu-Ghazaleh, Nael
contents Large Language Models (LLMs) are aligned to meet ethical standards and safety requirements by training them to refuse answering harmful or unsafe prompts. In this paper, we demonstrate how adversaries can exploit LLMs' alignment to implant bias, or enforce targeted censorship without degrading the model's responsiveness to unrelated topics. Specifically, we propose Subversive Alignment Injection (SAI), a poisoning attack that leverages the alignment mechanism to trigger refusal on specific topics or queries predefined by the adversary. Although it is perhaps not surprising that refusal can be induced through overalignment, we demonstrate how this refusal can be exploited to inject bias into the model. Surprisingly, SAI evades state-of-the-art poisoning defenses including LLM state forensics, as well as robust aggregation techniques that are designed to detect poisoning in FL settings. We demonstrate the practical dangers of this attack by illustrating its end-to-end impacts on LLM-powered application pipelines. For chat based applications such as ChatDoctor, with 1% data poisoning, the system refuses to answer healthcare questions to targeted racial category leading to high bias ($ΔDP$ of 23%). We also show that bias can be induced in other NLP tasks: for a resume selection pipeline aligned to refuse to summarize CVs from a selected university, high bias in selection ($ΔDP$ of 27%) results. Even higher bias ($ΔDP$~38%) results on 9 other chat based downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
Mamun, Md Abdullah Al
Alouani, Ihsen
Abu-Ghazaleh, Nael
Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
Large Language Models (LLMs) are aligned to meet ethical standards and safety requirements by training them to refuse answering harmful or unsafe prompts. In this paper, we demonstrate how adversaries can exploit LLMs' alignment to implant bias, or enforce targeted censorship without degrading the model's responsiveness to unrelated topics. Specifically, we propose Subversive Alignment Injection (SAI), a poisoning attack that leverages the alignment mechanism to trigger refusal on specific topics or queries predefined by the adversary. Although it is perhaps not surprising that refusal can be induced through overalignment, we demonstrate how this refusal can be exploited to inject bias into the model. Surprisingly, SAI evades state-of-the-art poisoning defenses including LLM state forensics, as well as robust aggregation techniques that are designed to detect poisoning in FL settings. We demonstrate the practical dangers of this attack by illustrating its end-to-end impacts on LLM-powered application pipelines. For chat based applications such as ChatDoctor, with 1% data poisoning, the system refuses to answer healthcare questions to targeted racial category leading to high bias ($ΔDP$ of 23%). We also show that bias can be induced in other NLP tasks: for a resume selection pipeline aligned to refuse to summarize CVs from a selected university, high bias in selection ($ΔDP$ of 27%) results. Even higher bias ($ΔDP$~38%) results on 9 other chat based downstream applications.
title Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.20333