Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
Fuente:
arXiv
Saved in:
| Main Authors: | Vyas, Sanyam, Hicks, Chris, Mavroudis, Vasilios |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025)
by: Vyas, Sanyam, et al.
Published: (2025)
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024)
by: Foley, Myles, et al.
Published: (2024)
Less is more? Rewards in RL for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2025)
by: Bates, Elizabeth, et al.
Published: (2025)
SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity
by: McFadden, Shae, et al.
Published: (2026)
by: McFadden, Shae, et al.
Published: (2026)
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024)
by: Thompson, Isaac Symes, et al.
Published: (2024)
An Attentive Graph Agent for Topology-Adaptive Cyber Defence
by: Sandoval, Ilya Orson, et al.
Published: (2025)
by: Sandoval, Ilya Orson, et al.
Published: (2025)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
DRMD: Deep Reinforcement Learning for Malware Detection under Concept Drift
by: McFadden, Shae, et al.
Published: (2025)
by: McFadden, Shae, et al.
Published: (2025)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Referential Security as a New Paradigm for AI Evaluations
by: Ristea, Dan, et al.
Published: (2026)
by: Ristea, Dan, et al.
Published: (2026)
Beyond Rewards in Reinforcement Learning for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2026)
by: Bates, Elizabeth, et al.
Published: (2026)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2026)
by: Ma, Oubo, et al.
Published: (2026)
UNIDOOR: A Universal Framework for Action-Level Backdoor Attacks in Deep Reinforcement Learning
by: Ma, Oubo, et al.
Published: (2025)
by: Ma, Oubo, et al.
Published: (2025)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Kill it with FIRE: On Leveraging Latent Space Directions for Runtime Backdoor Mitigation in Deep Neural Networks
by: Ahlers, Enrico, et al.
Published: (2026)
by: Ahlers, Enrico, et al.
Published: (2026)
TEN-GUARD: Tensor Decomposition for Backdoor Attack Detection in Deep Neural Networks
by: Hossain, Khondoker Murad, et al.
Published: (2024)
by: Hossain, Khondoker Murad, et al.
Published: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems
by: Fang, Jing, et al.
Published: (2025)
by: Fang, Jing, et al.
Published: (2025)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
Backdoor Attack on Vertical Federated Graph Neural Network Learning
by: Yang, Jirui, et al.
Published: (2024)
by: Yang, Jirui, et al.
Published: (2024)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
ReVeil: Unconstrained Concealed Backdoor Attack on Deep Neural Networks using Machine Unlearning
by: Alam, Manaar, et al.
Published: (2025)
by: Alam, Manaar, et al.
Published: (2025)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
by: Yang, Xiao, et al.
Published: (2024)
by: Yang, Xiao, et al.
Published: (2024)
Flatness-aware Sequential Learning Generates Resilient Backdoors
by: Pham, Hoang, et al.
Published: (2024)
by: Pham, Hoang, et al.
Published: (2024)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
by: Jian, Wang, et al.
Published: (2026)
by: Jian, Wang, et al.
Published: (2026)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
Differentially Private Deep Model-Based Reinforcement Learning
by: Rio, Alexandre, et al.
Published: (2024)
by: Rio, Alexandre, et al.
Published: (2024)
Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection
by: Høyheim, Eirik, et al.
Published: (2026)
by: Høyheim, Eirik, et al.
Published: (2026)
Client-Side Patching against Backdoor Attacks in Federated Learning
by: Molina-Coronado, Borja
Published: (2024)
by: Molina-Coronado, Borja
Published: (2024)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024)
by: Wu, Jiahao, et al.
Published: (2024)
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
by: Etter, Brian, et al.
Published: (2024)
by: Etter, Brian, et al.
Published: (2024)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
by: Pal, Soumyadeep, et al.
Published: (2024)
by: Pal, Soumyadeep, et al.
Published: (2024)
FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in FL
by: Hallaji, Ehsan, et al.
Published: (2025)
by: Hallaji, Ehsan, et al.
Published: (2025)
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Heterogeneous Graph Backdoor Attack
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Similar Items
-
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025) -
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024) -
Less is more? Rewards in RL for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2025) -
SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity
by: McFadden, Shae, et al.
Published: (2026) -
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024)