From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Guangyu, Cheng, Siyuan, Xu, Xiangzhe, Zhou, Yuan, Guo, Hanxi, Zhang, Zhuo, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
by: Yan, Lu, et al.
Published: (2025)
by: Yan, Lu, et al.
Published: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
by: Yan, Lu, et al.
Published: (2024)
by: Yan, Lu, et al.
Published: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
by: Wan, Wei, et al.
Published: (2025)
by: Wan, Wei, et al.
Published: (2025)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
The Stronger the Diffusion Model, the Easier the Backdoor: Data Poisoning to Induce Copyright Breaches Without Adjusting Finetuning Pipeline
by: Wang, Haonan, et al.
Published: (2024)
by: Wang, Haonan, et al.
Published: (2024)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
by: Wang, Xuan, et al.
Published: (2025)
by: Wang, Xuan, et al.
Published: (2025)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
by: Tie, Guiyao, et al.
Published: (2026)
by: Tie, Guiyao, et al.
Published: (2026)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Toward Polymorphic Backdoor against Semantic Communication via Intensity-Based Poisoning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
CBPF: Filtering Poisoned Data Based on Composite Backdoor Attack
by: Xia, Hanfeng, et al.
Published: (2024)
by: Xia, Hanfeng, et al.
Published: (2024)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
by: Jian, Wang, et al.
Published: (2026)
by: Jian, Wang, et al.
Published: (2026)
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Does Few-shot Learning Suffer from Backdoor Attacks?
by: Liu, Xinwei, et al.
Published: (2023)
by: Liu, Xinwei, et al.
Published: (2023)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
On The Dangers of Poisoned LLMs In Security Automation
by: Karlsen, Patrick, et al.
Published: (2025)
by: Karlsen, Patrick, et al.
Published: (2025)
Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning
by: Zhang, Yujie, et al.
Published: (2024)
by: Zhang, Yujie, et al.
Published: (2024)
Class-Conditional Neural Polarizer: A Lightweight and Effective Backdoor Defense by Purifying Poisoned Features
by: Zhu, Mingli, et al.
Published: (2025)
by: Zhu, Mingli, et al.
Published: (2025)
A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
by: Wu, Zhixiao, et al.
Published: (2025)
by: Wu, Zhixiao, et al.
Published: (2025)
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024)
by: Liu, Aishan, et al.
Published: (2024)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
by: Yang, Feiyu, et al.
Published: (2025)
by: Yang, Feiyu, et al.
Published: (2025)
TASER: Task-Aware Spectral Energy Refine for Backdoor Suppression in UAV Swarms Decentralized Federated Learning
by: Huang, Sizhe, et al.
Published: (2026)
by: Huang, Sizhe, et al.
Published: (2026)
Opening A Pandora's Box: Things You Should Know in the Era of Custom GPTs
by: Tao, Guanhong, et al.
Published: (2023)
by: Tao, Guanhong, et al.
Published: (2023)
Concept-Guided Backdoor Attack on Vision Language Models
by: Shen, Haoyu, et al.
Published: (2025)
by: Shen, Haoyu, et al.
Published: (2025)
CSC: Turning the Adversary's Poison against Itself
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
by: Zhou, Yihe, et al.
Published: (2025)
by: Zhou, Yihe, et al.
Published: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
by: Wang, Bingzheng, et al.
Published: (2026)
by: Wang, Bingzheng, et al.
Published: (2026)
SAB:A Stealing and Robust Backdoor Attack based on Steganographic Algorithm against Federated Learning
by: Xu, Weida, et al.
Published: (2024)
by: Xu, Weida, et al.
Published: (2024)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs
by: Yifeng, Peng, et al.
Published: (2025)
by: Yifeng, Peng, et al.
Published: (2025)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
by: Fan, Kaisheng, et al.
Published: (2026)
by: Fan, Kaisheng, et al.
Published: (2026)
Similar Items
-
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
by: Yan, Lu, et al.
Published: (2025) -
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
by: Guo, Weiyang, et al.
Published: (2026) -
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
by: Yan, Lu, et al.
Published: (2024) -
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024) -
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
by: Wan, Wei, et al.
Published: (2025)