BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruyi, Gao, Heng, Jian, Songlei, Tan, Yusong, Zhou, Haifang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914399437979648
author Zhang, Ruyi
Gao, Heng
Jian, Songlei
Tan, Yusong
Zhou, Haifang
author_facet Zhang, Ruyi
Gao, Heng
Jian, Songlei
Tan, Yusong
Zhou, Haifang
contents Backdoor attacks compromise model reliability by using triggers to manipulate outputs. Trigger inversion can accurately locate these triggers via a generator and is therefore critical for backdoor defense. However, the discrete nature of text prevents existing noise-based trigger generator from being applied to nature language processing (NLP). To overcome the limitations, we employ the rich knowledge embedded in large language models (LLMs) and propose a Backdoor defender powered by LLM Trigger Generator, termed BadLLM-TG. It is optimized through prompt-driven reinforcement learning, using the victim model's feedback loss as the reward signal. The generated triggers are then employed to mitigate the backdoor via adversarial training. Experiments show that our method reduces the attack success rate by 76.2\% on average, outperforming the second-best defender by 13.7.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15692
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
Zhang, Ruyi
Gao, Heng
Jian, Songlei
Tan, Yusong
Zhou, Haifang
Cryptography and Security
Artificial Intelligence
Backdoor attacks compromise model reliability by using triggers to manipulate outputs. Trigger inversion can accurately locate these triggers via a generator and is therefore critical for backdoor defense. However, the discrete nature of text prevents existing noise-based trigger generator from being applied to nature language processing (NLP). To overcome the limitations, we employ the rich knowledge embedded in large language models (LLMs) and propose a Backdoor defender powered by LLM Trigger Generator, termed BadLLM-TG. It is optimized through prompt-driven reinforcement learning, using the victim model's feedback loss as the reward signal. The generated triggers are then employed to mitigate the backdoor via adversarial training. Experiments show that our method reduces the attack success rate by 76.2\% on average, outperforming the second-best defender by 13.7.
title BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.15692