Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
Fuente:
arXiv
Saved in:
| Main Authors: | Braun, Tobias, Grebe, Jonas Henry, Rohrbach, Marcus, Rohrbach, Anna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
by: Braun, Tobias, et al.
Published: (2026)
by: Braun, Tobias, et al.
Published: (2026)
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026)
by: Grebe, Jonas Henry, et al.
Published: (2026)
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024)
by: Liu, Aishan, et al.
Published: (2024)
Rethinking the Vulnerability of Concept Erasure and a New Method
by: Richardson, Alex D., et al.
Published: (2025)
by: Richardson, Alex D., et al.
Published: (2025)
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
by: Jia, Hengrui, et al.
Published: (2025)
by: Jia, Hengrui, et al.
Published: (2025)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Erasing Self-Supervised Learning Backdoor by Cluster Activation Masking
by: Qian, Shengsheng, et al.
Published: (2023)
by: Qian, Shengsheng, et al.
Published: (2023)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy
by: Sui, Zhihao, et al.
Published: (2025)
by: Sui, Zhihao, et al.
Published: (2025)
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024)
by: Wu, Jiahao, et al.
Published: (2024)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
by: Pal, Soumyadeep, et al.
Published: (2024)
by: Pal, Soumyadeep, et al.
Published: (2024)
Heterogeneous Graph Backdoor Attack
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
Mirror Mirror on the Wall, Have I Forgotten it All? A New Framework for Evaluating Machine Unlearning
by: Brimhall, Brennon, et al.
Published: (2025)
by: Brimhall, Brennon, et al.
Published: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
Backdoor Detection through Replicated Execution of Outsourced Training
by: Jia, Hengrui, et al.
Published: (2025)
by: Jia, Hengrui, et al.
Published: (2025)
PBP: Post-training Backdoor Purification for Malware Classifiers
by: Nguyen, Dung Thuy, et al.
Published: (2024)
by: Nguyen, Dung Thuy, et al.
Published: (2024)
On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
by: Pichler, Georg, et al.
Published: (2024)
by: Pichler, Georg, et al.
Published: (2024)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
by: Jian, Wang, et al.
Published: (2026)
by: Jian, Wang, et al.
Published: (2026)
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Flatness-aware Sequential Learning Generates Resilient Backdoors
by: Pham, Hoang, et al.
Published: (2024)
by: Pham, Hoang, et al.
Published: (2024)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
Your Agent Can Defend Itself against Backdoor Attacks
by: Changjiang, Li, et al.
Published: (2025)
by: Changjiang, Li, et al.
Published: (2025)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
Revisiting Backdoor Attacks on Time Series Classification in the Frequency Domain
by: Huang, Yuanmin, et al.
Published: (2025)
by: Huang, Yuanmin, et al.
Published: (2025)
Fast and Lightweight Backdoor Detection via Head Random Probing
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
by: Yang, Xiao, et al.
Published: (2024)
by: Yang, Xiao, et al.
Published: (2024)
Similar Items
-
Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models
by: Braun, Tobias, et al.
Published: (2026) -
Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
by: Grebe, Jonas Henry, et al.
Published: (2026) -
Compromising Embodied Agents with Contextual Backdoor Attacks
by: Liu, Aishan, et al.
Published: (2024) -
Rethinking the Vulnerability of Concept Erasure and a New Method
by: Richardson, Alex D., et al.
Published: (2025) -
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
by: Jia, Hengrui, et al.
Published: (2025)