Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jianwei, Kim, Jung-Eun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
von: Li, Jianwei, et al.
Veröffentlicht: (2026)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
von: Li, Jianwei, et al.
Veröffentlicht: (2025)
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
von: Li, Jianwei, et al.
Veröffentlicht: (2024)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
von: Chen, Depeng, et al.
Veröffentlicht: (2024)
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
von: Kirci, Onur Alp, et al.
Veröffentlicht: (2025)
von: Kirci, Onur Alp, et al.
Veröffentlicht: (2025)
A Semantic and Clean-label Backdoor Attack against Graph Convolutional Networks
von: Dai, Jiazhu, et al.
Veröffentlicht: (2025)
von: Dai, Jiazhu, et al.
Veröffentlicht: (2025)
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
How to Backdoor the Knowledge Distillation
von: Wu, Chen, et al.
Veröffentlicht: (2025)
von: Wu, Chen, et al.
Veröffentlicht: (2025)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
Decoupling Generalizability and Membership Privacy Risks in Neural Networks
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
von: Fang, Xingli, et al.
Veröffentlicht: (2026)
Center-Based Relaxed Learning Against Membership Inference Attacks
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
von: Fang, Xingli, et al.
Veröffentlicht: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Backdoor Graph Condensation
von: Wu, Jiahao, et al.
Veröffentlicht: (2024)
von: Wu, Jiahao, et al.
Veröffentlicht: (2024)
Flatness-aware Sequential Learning Generates Resilient Backdoors
von: Pham, Hoang, et al.
Veröffentlicht: (2024)
von: Pham, Hoang, et al.
Veröffentlicht: (2024)
Heterogeneous Graph Backdoor Attack
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
von: Lu, Junchi, et al.
Veröffentlicht: (2025)
von: Lu, Junchi, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
von: Pawlak, Stanisław, et al.
Veröffentlicht: (2025)
von: Pawlak, Stanisław, et al.
Veröffentlicht: (2025)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
von: Li, Nan, et al.
Veröffentlicht: (2024)
von: Li, Nan, et al.
Veröffentlicht: (2024)
Magnitude-based Neuron Pruning for Backdoor Defens
von: Li, Nan, et al.
Veröffentlicht: (2024)
von: Li, Nan, et al.
Veröffentlicht: (2024)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
von: Shin, Jeongjin, et al.
Veröffentlicht: (2024)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
von: Yang, Xiao, et al.
Veröffentlicht: (2024)
von: Yang, Xiao, et al.
Veröffentlicht: (2024)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Backdoor defense, learnability and obfuscation
von: Christiano, Paul, et al.
Veröffentlicht: (2024)
von: Christiano, Paul, et al.
Veröffentlicht: (2024)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
von: Min, Rui, et al.
Veröffentlicht: (2024)
von: Min, Rui, et al.
Veröffentlicht: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
von: Mo, Yichuan, et al.
Veröffentlicht: (2024)
Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
von: Yi, Bongsoo, et al.
Veröffentlicht: (2024)
Towards Sample-specific Backdoor Attack with Clean Labels via Attribute Trigger
von: Zhu, Mingyan, et al.
Veröffentlicht: (2023)
von: Zhu, Mingyan, et al.
Veröffentlicht: (2023)
Your Agent Can Defend Itself against Backdoor Attacks
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
von: Changjiang, Li, et al.
Veröffentlicht: (2025)
Revisiting Backdoor Attacks on Time Series Classification in the Frequency Domain
von: Huang, Yuanmin, et al.
Veröffentlicht: (2025)
von: Huang, Yuanmin, et al.
Veröffentlicht: (2025)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
von: Gong, Chen, et al.
Veröffentlicht: (2022)
von: Gong, Chen, et al.
Veröffentlicht: (2022)
OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection
von: Jiang, Haoyu, et al.
Veröffentlicht: (2023)
von: Jiang, Haoyu, et al.
Veröffentlicht: (2023)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
von: Han, Tingxu, et al.
Veröffentlicht: (2024)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
How to Craft Backdoors with Unlabeled Data Alone?
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation
von: Li, Jianwei, et al.
Veröffentlicht: (2026) -
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
von: Li, Jianwei, et al.
Veröffentlicht: (2025) -
Superficial Safety Alignment Hypothesis
von: Li, Jianwei, et al.
Veröffentlicht: (2024) -
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
von: Chen, Depeng, et al.
Veröffentlicht: (2024) -
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
von: Kirci, Onur Alp, et al.
Veröffentlicht: (2025)