The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
Fuente:
arXiv
Saved in:
| Main Authors: | Bullwinkel, Blake, Severi, Giorgio, Hines, Keegan, Minnich, Amanda, Kumar, Ram Shankar Siva, Zunger, Yonatan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
by: Bullwinkel, Blake, et al.
Published: (2025)
by: Bullwinkel, Blake, et al.
Published: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
by: Yang, Feiyu, et al.
Published: (2025)
by: Yang, Feiyu, et al.
Published: (2025)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
by: Ma, Yuan, et al.
Published: (2024)
by: Ma, Yuan, et al.
Published: (2024)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
by: Chen, Shenyang, et al.
Published: (2026)
by: Chen, Shenyang, et al.
Published: (2026)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
by: Xue, Xiaoyu, et al.
Published: (2025)
by: Xue, Xiaoyu, et al.
Published: (2025)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
by: Fan, Kaisheng, et al.
Published: (2026)
by: Fan, Kaisheng, et al.
Published: (2026)
A Systematization of Security Vulnerabilities in Computer Use Agents
by: Jones, Daniel, et al.
Published: (2025)
by: Jones, Daniel, et al.
Published: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
The Art of Deception: Robust Backdoor Attack using Dynamic Stacking of Triggers
by: Mengara, Orson
Published: (2024)
by: Mengara, Orson
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning
by: Zhang, Yujie, et al.
Published: (2024)
by: Zhang, Yujie, et al.
Published: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
by: Wu, Zhixiao, et al.
Published: (2025)
by: Wu, Zhixiao, et al.
Published: (2025)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
by: Pang, Yan, et al.
Published: (2025)
by: Pang, Yan, et al.
Published: (2025)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Re-Triggering Safeguards within LLMs for Jailbreak Detection
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Discovering Universal Semantic Triggers for Text-to-Image Synthesis
by: Zhai, Shengfang, et al.
Published: (2024)
by: Zhai, Shengfang, et al.
Published: (2024)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
by: Russinovich, Mark, et al.
Published: (2026)
by: Russinovich, Mark, et al.
Published: (2026)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
ShadowLogic: Backdoors in Any Whitebox LLM
by: Schulz, Kasimir, et al.
Published: (2025)
by: Schulz, Kasimir, et al.
Published: (2025)
RESTRAIN: Reinforcement Learning-Based Secure Framework for Trigger-Action IoT Environment
by: Alam, Md Morshed, et al.
Published: (2025)
by: Alam, Md Morshed, et al.
Published: (2025)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
by: Sun, Zehan, et al.
Published: (2026)
by: Sun, Zehan, et al.
Published: (2026)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
by: Ding, Renhua, et al.
Published: (2025)
by: Ding, Renhua, et al.
Published: (2025)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
by: Ning, Liangbo, et al.
Published: (2025)
by: Ning, Liangbo, et al.
Published: (2025)
Collaborative Intelligence: Topic Modelling of Large Language Model use in Live Cybersecurity Operations
by: Lochner, Martin, et al.
Published: (2025)
by: Lochner, Martin, et al.
Published: (2025)
Towards Sample-specific Backdoor Attack with Clean Labels via Attribute Trigger
by: Zhu, Mingyan, et al.
Published: (2023)
by: Zhu, Mingyan, et al.
Published: (2023)
Similar Items
-
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
by: Bullwinkel, Blake, et al.
Published: (2025) -
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026) -
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024) -
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025) -
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)