Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Qi, Ye, Zipeng, Tang, Yubo, Luo, Wenjian, Shi, Yuhui, Jia, Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lightweight and Fast Backdoor Model Detection
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
von: Fan, Kaisheng, et al.
Veröffentlicht: (2026)
von: Fan, Kaisheng, et al.
Veröffentlicht: (2026)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
Class-Conditional Neural Polarizer: A Lightweight and Effective Backdoor Defense by Purifying Poisoned Features
von: Zhu, Mingli, et al.
Veröffentlicht: (2025)
von: Zhu, Mingli, et al.
Veröffentlicht: (2025)
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2026)
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2026)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
von: Wan, Wei, et al.
Veröffentlicht: (2025)
von: Wan, Wei, et al.
Veröffentlicht: (2025)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Fast and Lightweight Backdoor Detection via Head Random Probing
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
von: Xue, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xue, Xiaoyu, et al.
Veröffentlicht: (2025)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
von: Ning, Liangbo, et al.
Veröffentlicht: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
von: Zhang, Ruyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruyi, et al.
Veröffentlicht: (2026)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
von: Lin, Chenhao, et al.
Veröffentlicht: (2025)
von: Lin, Chenhao, et al.
Veröffentlicht: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
von: Hou, Yang, et al.
Veröffentlicht: (2024)
von: Hou, Yang, et al.
Veröffentlicht: (2024)
A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers
von: Wu, Zhixiao, et al.
Veröffentlicht: (2025)
von: Wu, Zhixiao, et al.
Veröffentlicht: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
von: Qiu, Huming, et al.
Veröffentlicht: (2023)
von: Qiu, Huming, et al.
Veröffentlicht: (2023)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
von: Wang, Qingyue, et al.
Veröffentlicht: (2025)
von: Wang, Qingyue, et al.
Veröffentlicht: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
von: Chen, Shenyang, et al.
Veröffentlicht: (2026)
von: Chen, Shenyang, et al.
Veröffentlicht: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
von: Yan, Shenao, et al.
Veröffentlicht: (2024)
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
von: Tie, Guiyao, et al.
Veröffentlicht: (2026)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
von: Pawlak, Stanisław, et al.
Veröffentlicht: (2025)
von: Pawlak, Stanisław, et al.
Veröffentlicht: (2025)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
von: Popovic, Dorde, et al.
Veröffentlicht: (2025)
von: Popovic, Dorde, et al.
Veröffentlicht: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
von: Peng, Zuquan, et al.
Veröffentlicht: (2025)
von: Peng, Zuquan, et al.
Veröffentlicht: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
von: Min, Rui, et al.
Veröffentlicht: (2024)
von: Min, Rui, et al.
Veröffentlicht: (2024)
PAD-FT: A Lightweight Defense for Backdoor Attacks via Data Purification and Fine-Tuning
von: Xu, Yukai, et al.
Veröffentlicht: (2024)
von: Xu, Yukai, et al.
Veröffentlicht: (2024)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lightweight and Fast Backdoor Model Detection
von: Yu, Yinbo, et al.
Veröffentlicht: (2026) -
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
von: Fan, Kaisheng, et al.
Veröffentlicht: (2026) -
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
von: Ma, Yuan, et al.
Veröffentlicht: (2024) -
Class-Conditional Neural Polarizer: A Lightweight and Effective Backdoor Defense by Purifying Poisoned Features
von: Zhu, Mingli, et al.
Veröffentlicht: (2025) -
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
von: Bullwinkel, Blake, et al.
Veröffentlicht: (2026)