BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ruyi, Gao, Heng, Jian, Songlei, Tan, Yusong, Zhou, Haifang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
by: Bullwinkel, Blake, et al.
Published: (2026)
by: Bullwinkel, Blake, et al.
Published: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
by: Pang, Yan, et al.
Published: (2025)
by: Pang, Yan, et al.
Published: (2025)
BadEdit: Backdooring large language models by model editing
by: Li, Yanzhou, et al.
Published: (2024)
by: Li, Yanzhou, et al.
Published: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
ShadowLogic: Backdoors in Any Whitebox LLM
by: Schulz, Kasimir, et al.
Published: (2025)
by: Schulz, Kasimir, et al.
Published: (2025)
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
by: Yao, Yifan, et al.
Published: (2023)
by: Yao, Yifan, et al.
Published: (2023)
DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent
by: Zhu, Pengyu, et al.
Published: (2025)
by: Zhu, Pengyu, et al.
Published: (2025)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
by: Zhong, Peter Yong, et al.
Published: (2025)
by: Zhong, Peter Yong, et al.
Published: (2025)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
by: Jin, Haotian, et al.
Published: (2025)
by: Jin, Haotian, et al.
Published: (2025)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
by: Lin, Junda, et al.
Published: (2026)
by: Lin, Junda, et al.
Published: (2026)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
by: Ning, Liangbo, et al.
Published: (2025)
by: Ning, Liangbo, et al.
Published: (2025)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
by: Fan, Kaisheng, et al.
Published: (2026)
by: Fan, Kaisheng, et al.
Published: (2026)
MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks
by: Zhou, Tong, et al.
Published: (2025)
by: Zhou, Tong, et al.
Published: (2025)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron
by: Miah, Abdullah Arafat, et al.
Published: (2026)
by: Miah, Abdullah Arafat, et al.
Published: (2026)
A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
by: Wu, Zhixiao, et al.
Published: (2025)
by: Wu, Zhixiao, et al.
Published: (2025)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
by: Xue, Xiaoyu, et al.
Published: (2025)
by: Xue, Xiaoyu, et al.
Published: (2025)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
by: Yang, Feiyu, et al.
Published: (2025)
by: Yang, Feiyu, et al.
Published: (2025)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
by: Ma, Yuan, et al.
Published: (2024)
by: Ma, Yuan, et al.
Published: (2024)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
by: Li, Zi, et al.
Published: (2026)
by: Li, Zi, et al.
Published: (2026)
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use
by: Zhang, Wuyang, et al.
Published: (2026)
by: Zhang, Wuyang, et al.
Published: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
Buffer is All You Need: Defending Federated Learning against Backdoor Attacks under Non-iids via Buffering
by: Lyu, Xingyu, et al.
Published: (2025)
by: Lyu, Xingyu, et al.
Published: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Similar Items
-
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
by: Bullwinkel, Blake, et al.
Published: (2026) -
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025) -
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
by: Pang, Yan, et al.
Published: (2025) -
BadEdit: Backdooring large language models by model editing
by: Li, Yanzhou, et al.
Published: (2024) -
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)