Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Kaisheng, Zhang, Weizhe, Gao, Yishu, Bissyandé, Tegawendé F., Tang, Xunzhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MalLoc: Toward Fine-grained Android Malicious Payload Localization via LLMs
von: Sun, Tiezhu, et al.
Veröffentlicht: (2025)
von: Sun, Tiezhu, et al.
Veröffentlicht: (2025)
Assessing Spear-Phishing Website Generation in Large Language Model Coding Agents
von: Malloy, Tailia, et al.
Veröffentlicht: (2026)
von: Malloy, Tailia, et al.
Veröffentlicht: (2026)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
von: Tessa, Melissa, et al.
Veröffentlicht: (2026)
von: Tessa, Melissa, et al.
Veröffentlicht: (2026)
(In)Security of Mobile Apps in Developing Countries: A Systematic Literature Review
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in detecting Secrets in Android Apps
von: Alecci, Marco, et al.
Veröffentlicht: (2025)
von: Alecci, Marco, et al.
Veröffentlicht: (2025)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
ProxyPrints: From Database Breach to Spoof, A Plug-and-Play Defense for Biometric Systems
von: Hacmon, Yaniv, et al.
Veröffentlicht: (2025)
von: Hacmon, Yaniv, et al.
Veröffentlicht: (2025)
Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks
von: Fu, Yanzhang, et al.
Veröffentlicht: (2026)
von: Fu, Yanzhang, et al.
Veröffentlicht: (2026)
Software Security in Software-Defined Networking: A Systematic Literature Review
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
von: Diouf, Moustapha Awwalou, et al.
Veröffentlicht: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
Backdoor Attacks and Defenses in Computer Vision Domain: A Survey
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
Backdoor Contrastive Learning via Bi-level Trigger Optimization
von: Sun, Weiyu, et al.
Veröffentlicht: (2024)
von: Sun, Weiyu, et al.
Veröffentlicht: (2024)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
Security Assessment of Mobile Banking Apps in West African Economic and Monetary Union
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
von: Diallo, Alioune, et al.
Veröffentlicht: (2024)
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
von: Lin, Zheng, et al.
Veröffentlicht: (2026)
BDPFL: Backdoor Defense for Personalized Federated Learning via Explainable Distillation
von: Zhu, Chengcheng, et al.
Veröffentlicht: (2025)
von: Zhu, Chengcheng, et al.
Veröffentlicht: (2025)
Attack as Defense: Run-time Backdoor Implantation for Image Content Protection
von: Zhang, Haichuan, et al.
Veröffentlicht: (2024)
von: Zhang, Haichuan, et al.
Veröffentlicht: (2024)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
Backdoor Threats in Variational Quantum Circuits: Taxonomy, Attacks, and Defenses
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
Hardware-Triggered Backdoors
von: Möller, Jonas, et al.
Veröffentlicht: (2026)
von: Möller, Jonas, et al.
Veröffentlicht: (2026)
ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning
von: Shui, Yuhan, et al.
Veröffentlicht: (2026)
von: Shui, Yuhan, et al.
Veröffentlicht: (2026)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
von: Zhao, Xin, et al.
Veröffentlicht: (2025)
von: Zhao, Xin, et al.
Veröffentlicht: (2025)
DexRay: A Simple, yet Effective Deep Learning Approach to Android Malware Detection based on Image Representation of Bytecode
von: Daoudi, Nadia, et al.
Veröffentlicht: (2021)
von: Daoudi, Nadia, et al.
Veröffentlicht: (2021)
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
von: Abad, Gorka, et al.
Veröffentlicht: (2026)
von: Abad, Gorka, et al.
Veröffentlicht: (2026)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2026)
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2026)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
von: Zhai, Shengfang, et al.
Veröffentlicht: (2025)
von: Zhai, Shengfang, et al.
Veröffentlicht: (2025)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
A Practical Trigger-Free Backdoor Attack on Neural Networks
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
Universal Graph Backdoor Defense: A Feature-based Homophily Perspective
von: Pan, Mengting, et al.
Veröffentlicht: (2026)
von: Pan, Mengting, et al.
Veröffentlicht: (2026)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MalLoc: Toward Fine-grained Android Malicious Payload Localization via LLMs
von: Sun, Tiezhu, et al.
Veröffentlicht: (2025) -
Assessing Spear-Phishing Website Generation in Large Language Model Coding Agents
von: Malloy, Tailia, et al.
Veröffentlicht: (2026) -
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
von: Tessa, Melissa, et al.
Veröffentlicht: (2026) -
(In)Security of Mobile Apps in Developing Countries: A Systematic Literature Review
von: Diallo, Alioune, et al.
Veröffentlicht: (2024) -
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
von: Zhou, Qi, et al.
Veröffentlicht: (2024)