Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Pallakonda, Bhanu, Hindsbo, Mikkel, Ehsani, Sina, Mishra, Prag |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Injecting Bias into Text Classification Models using Backdoor Attacks
by: Yavuz, A. Dilara, et al.
Published: (2024)
by: Yavuz, A. Dilara, et al.
Published: (2024)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025)
by: Boisvert, Léo, et al.
Published: (2025)
Camera Control at the Edge with Language Models for Scene Understanding
by: Buynitsky, Alexiy, et al.
Published: (2025)
by: Buynitsky, Alexiy, et al.
Published: (2025)
Prompt Inject Detection with Generative Explanation as an Investigative Tool
by: Pan, Jonathan, et al.
Published: (2025)
by: Pan, Jonathan, et al.
Published: (2025)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
by: Yu, Haiyang, et al.
Published: (2024)
by: Yu, Haiyang, et al.
Published: (2024)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
by: Wang, Bingzheng, et al.
Published: (2026)
by: Wang, Bingzheng, et al.
Published: (2026)
Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
by: D'addario, Andrew Maranhão Ventura
Published: (2025)
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
by: Shen, Guangyu, et al.
Published: (2025)
by: Shen, Guangyu, et al.
Published: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
Beyond Immediate Activation: Temporally Decoupled Backdoor Attacks on Time Series Forecasting
by: Liu, Zhixin, et al.
Published: (2026)
by: Liu, Zhixin, et al.
Published: (2026)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
by: Zhou, Yihe, et al.
Published: (2025)
by: Zhou, Yihe, et al.
Published: (2025)
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use
by: Zhang, Wuyang, et al.
Published: (2026)
by: Zhang, Wuyang, et al.
Published: (2026)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
by: Fastowski, Alina, et al.
Published: (2025)
by: Fastowski, Alina, et al.
Published: (2025)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
by: Fan, Kaisheng, et al.
Published: (2026)
by: Fan, Kaisheng, et al.
Published: (2026)
Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
by: Liu, Fazhong, et al.
Published: (2026)
by: Liu, Fazhong, et al.
Published: (2026)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
by: Riaño, Roberto, et al.
Published: (2024)
by: Riaño, Roberto, et al.
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
by: Popovic, Dorde, et al.
Published: (2025)
by: Popovic, Dorde, et al.
Published: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
Backdooring Bias in Large Language Models
by: Das, Anudeep, et al.
Published: (2026)
by: Das, Anudeep, et al.
Published: (2026)
Lightweight and Fast Backdoor Model Detection
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools?
by: Quantrill, Jason, et al.
Published: (2026)
by: Quantrill, Jason, et al.
Published: (2026)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Similar Items
-
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025) -
Injecting Bias into Text Classification Models using Backdoor Attacks
by: Yavuz, A. Dilara, et al.
Published: (2024) -
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025) -
Camera Control at the Edge with Language Models for Scene Understanding
by: Buynitsky, Alexiy, et al.
Published: (2025) -
Prompt Inject Detection with Generative Explanation as an Investigative Tool
by: Pan, Jonathan, et al.
Published: (2025)