BadActs: A Universal Backdoor Defense in the Activation Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yi, Biao, Chen, Sishuo, Li, Yiming, Li, Tong, Zhang, Baolei, Liu, Zheli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
von: Yin, Rui, et al.
Veröffentlicht: (2026)
von: Yin, Rui, et al.
Veröffentlicht: (2026)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
Practical Framework for Privacy-Preserving and Byzantine-robust Federated Learning
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
Traceback of Poisoning Attacks to Retrieval-Augmented Generation
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution
von: Tong, Yao, et al.
Veröffentlicht: (2024)
von: Tong, Yao, et al.
Veröffentlicht: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
von: Du, Wei, et al.
Veröffentlicht: (2023)
von: Du, Wei, et al.
Veröffentlicht: (2023)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
Data-centric NLP Backdoor Defense from the Lens of Memorization
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
von: Afane, Mohamed, et al.
Veröffentlicht: (2025)
von: Afane, Mohamed, et al.
Veröffentlicht: (2025)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
von: Wen, Rui, et al.
Veröffentlicht: (2026)
von: Wen, Rui, et al.
Veröffentlicht: (2026)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
von: Lyu, Weimin, et al.
Veröffentlicht: (2024)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
von: Zhou, Qin, et al.
Veröffentlicht: (2025)
von: Zhou, Qin, et al.
Veröffentlicht: (2025)
Nearest is Not Dearest: Towards Practical Defense against Quantization-conditioned Backdoor Attacks
von: Li, Boheng, et al.
Veröffentlicht: (2024)
von: Li, Boheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025) -
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
von: Yi, Biao, et al.
Veröffentlicht: (2025) -
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025) -
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025) -
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)