Backdoor defense, learnability and obfuscation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Christiano, Paul, Hilton, Jacob, Lecomte, Victor, Xu, Mark |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Detecting new obfuscated malware variants: A lightweight and interpretable machine learning approach
par: Madamidola, Oladipo A., et autres
Publié: (2024)
par: Madamidola, Oladipo A., et autres
Publié: (2024)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
par: De Muri, Giovanni, et autres
Publié: (2025)
par: De Muri, Giovanni, et autres
Publié: (2025)
Invisible Backdoor Attack Through Singular Value Decomposition
par: Chen, Wenmin, et autres
Publié: (2024)
par: Chen, Wenmin, et autres
Publié: (2024)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
par: Pawlak, Stanisław, et autres
Publié: (2025)
par: Pawlak, Stanisław, et autres
Publié: (2025)
Backdoor Graph Condensation
par: Wu, Jiahao, et autres
Publié: (2024)
par: Wu, Jiahao, et autres
Publié: (2024)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
par: Shin, Jeongjin, et autres
Publié: (2024)
par: Shin, Jeongjin, et autres
Publié: (2024)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
par: Pal, Soumyadeep, et autres
Publié: (2024)
par: Pal, Soumyadeep, et autres
Publié: (2024)
How to Backdoor the Knowledge Distillation
par: Wu, Chen, et autres
Publié: (2025)
par: Wu, Chen, et autres
Publié: (2025)
Heterogeneous Graph Backdoor Attack
par: Chen, Jiawei, et autres
Publié: (2025)
par: Chen, Jiawei, et autres
Publié: (2025)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
par: Gong, Chen, et autres
Publié: (2022)
par: Gong, Chen, et autres
Publié: (2022)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
par: Yang, Fan, et autres
Publié: (2026)
par: Yang, Fan, et autres
Publié: (2026)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
par: Li, Nan, et autres
Publié: (2024)
par: Li, Nan, et autres
Publié: (2024)
How to Craft Backdoors with Unlabeled Data Alone?
par: Wang, Yifei, et autres
Publié: (2024)
par: Wang, Yifei, et autres
Publié: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
par: Liu, Aishan, et autres
Publié: (2024)
par: Liu, Aishan, et autres
Publié: (2024)
Magnitude-based Neuron Pruning for Backdoor Defens
par: Li, Nan, et autres
Publié: (2024)
par: Li, Nan, et autres
Publié: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
par: Chen, Zhuowei, et autres
Publié: (2025)
par: Chen, Zhuowei, et autres
Publié: (2025)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
par: An, Shengwei, et autres
Publié: (2023)
par: An, Shengwei, et autres
Publié: (2023)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
par: Kong, Cong, et autres
Publié: (2024)
par: Kong, Cong, et autres
Publié: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
par: Carlini, Nicholas, et autres
Publié: (2025)
par: Carlini, Nicholas, et autres
Publié: (2025)
PBP: Post-training Backdoor Purification for Malware Classifiers
par: Nguyen, Dung Thuy, et autres
Publié: (2024)
par: Nguyen, Dung Thuy, et autres
Publié: (2024)
On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
par: Pichler, Georg, et autres
Publié: (2024)
par: Pichler, Georg, et autres
Publié: (2024)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
par: Li, Wei, et autres
Publié: (2024)
par: Li, Wei, et autres
Publié: (2024)
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
par: Lin, Xiao, et autres
Publié: (2024)
par: Lin, Xiao, et autres
Publié: (2024)
Backdooring Bias ($B^2$) into Stable Diffusion Models
par: Naseh, Ali, et autres
Publié: (2024)
par: Naseh, Ali, et autres
Publié: (2024)
Flatness-aware Sequential Learning Generates Resilient Backdoors
par: Pham, Hoang, et autres
Publié: (2024)
par: Pham, Hoang, et autres
Publié: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
par: Min, Rui, et autres
Publié: (2024)
par: Min, Rui, et autres
Publié: (2024)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
par: Ren, Zhiyao, et autres
Publié: (2025)
par: Ren, Zhiyao, et autres
Publié: (2025)
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
par: Braun, Tobias, et autres
Publié: (2025)
par: Braun, Tobias, et autres
Publié: (2025)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
par: Jian, Wang, et autres
Publié: (2026)
par: Jian, Wang, et autres
Publié: (2026)
Backdoor Detection through Replicated Execution of Outsourced Training
par: Jia, Hengrui, et autres
Publié: (2025)
par: Jia, Hengrui, et autres
Publié: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
par: Xu, Jiashu, et autres
Publié: (2023)
par: Xu, Jiashu, et autres
Publié: (2023)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
par: Yang, Xiao, et autres
Publié: (2024)
par: Yang, Xiao, et autres
Publié: (2024)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
par: Vyas, Sanyam, et autres
Publié: (2024)
par: Vyas, Sanyam, et autres
Publié: (2024)
Backdoor Attack on Vertical Federated Graph Neural Network Learning
par: Yang, Jirui, et autres
Publié: (2024)
par: Yang, Jirui, et autres
Publié: (2024)
Client-Side Patching against Backdoor Attacks in Federated Learning
par: Molina-Coronado, Borja
Publié: (2024)
par: Molina-Coronado, Borja
Publié: (2024)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
par: Han, Tingxu, et autres
Publié: (2024)
par: Han, Tingxu, et autres
Publié: (2024)
Your Agent Can Defend Itself against Backdoor Attacks
par: Changjiang, Li, et autres
Publié: (2025)
par: Changjiang, Li, et autres
Publié: (2025)
Fast and Lightweight Backdoor Detection via Head Random Probing
par: Yu, Yinbo, et autres
Publié: (2026)
par: Yu, Yinbo, et autres
Publié: (2026)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
par: Chen, Xiangxiang, et autres
Publié: (2025)
par: Chen, Xiangxiang, et autres
Publié: (2025)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
par: Foroughi, Mohammad Hadi, et autres
Publié: (2026)
par: Foroughi, Mohammad Hadi, et autres
Publié: (2026)
Documents similaires
-
Detecting new obfuscated malware variants: A lightweight and interpretable machine learning approach
par: Madamidola, Oladipo A., et autres
Publié: (2024) -
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
par: De Muri, Giovanni, et autres
Publié: (2025) -
Invisible Backdoor Attack Through Singular Value Decomposition
par: Chen, Wenmin, et autres
Publié: (2024) -
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
par: Pawlak, Stanisław, et autres
Publié: (2025) -
Backdoor Graph Condensation
par: Wu, Jiahao, et autres
Publié: (2024)