Robust Backdoor Removal by Reconstructing Trigger-Activated Changes in Latent Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Iwahana, Kazuki, Yamasaki, Yusuke, Ito, Akira, Miura, Takayuki, Shibahara, Toshiki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?
by: Ito, Akira, et al.
Published: (2025)
by: Ito, Akira, et al.
Published: (2025)
A Channel-Triggered Backdoor Attack on Wireless Semantic Image Reconstruction
by: Wan, Jialin, et al.
Published: (2025)
by: Wan, Jialin, et al.
Published: (2025)
Hardware-Triggered Backdoors
by: Möller, Jonas, et al.
Published: (2026)
by: Möller, Jonas, et al.
Published: (2026)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026)
by: Abad, Gorka, et al.
Published: (2026)
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
by: Ashcraft, Chace, et al.
Published: (2025)
by: Ashcraft, Chace, et al.
Published: (2025)
Cross-Paradigm Graph Backdoor Attacks with Promptable Subgraph Triggers
by: Liu, Dongyi, et al.
Published: (2025)
by: Liu, Dongyi, et al.
Published: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
The Art of Deception: Robust Backdoor Attack using Dynamic Stacking of Triggers
by: Mengara, Orson
Published: (2024)
by: Mengara, Orson
Published: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
BAN: Detecting Backdoors Activated by Adversarial Neuron Noise
by: Xu, Xiaoyun, et al.
Published: (2024)
by: Xu, Xiaoyun, et al.
Published: (2024)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
by: Guo, Zhen, et al.
Published: (2025)
by: Guo, Zhen, et al.
Published: (2025)
Robustness Inspired Graph Backdoor Defense
by: Zhang, Zhiwei, et al.
Published: (2024)
by: Zhang, Zhiwei, et al.
Published: (2024)
Is the Trigger Essential? A Feature-Based Triggerless Backdoor Attack in Vertical Federated Learning
by: Liu, Yige, et al.
Published: (2026)
by: Liu, Yige, et al.
Published: (2026)
On the Robustness of Graph Reduction Against GNN Backdoor
by: Zhu, Yuxuan, et al.
Published: (2024)
by: Zhu, Yuxuan, et al.
Published: (2024)
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
by: Zhang, Mingxuan, et al.
Published: (2025)
by: Zhang, Mingxuan, et al.
Published: (2025)
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
by: Eggen, Marte, et al.
Published: (2026)
by: Eggen, Marte, et al.
Published: (2026)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
Oblivious Defense in ML Models: Backdoor Removal without Detection
by: Goldwasser, Shafi, et al.
Published: (2024)
by: Goldwasser, Shafi, et al.
Published: (2024)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
by: Gosch, Lukas, et al.
Published: (2024)
by: Gosch, Lukas, et al.
Published: (2024)
Backdoor Attack with Sparse and Invisible Trigger
by: Gao, Yinghua, et al.
Published: (2023)
by: Gao, Yinghua, et al.
Published: (2023)
Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
by: Wu, Baoyuan, et al.
Published: (2024)
by: Wu, Baoyuan, et al.
Published: (2024)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026)
by: Saxena, Aman, et al.
Published: (2026)
Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning
by: Zhang, Yujie, et al.
Published: (2024)
by: Zhang, Yujie, et al.
Published: (2024)
Robustness bounds on the successful adversarial examples in probabilistic models: Implications from Gaussian processes
by: Maeshima, Hiroaki, et al.
Published: (2024)
by: Maeshima, Hiroaki, et al.
Published: (2024)
Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions
by: Cinà, Antonio Emanuele, et al.
Published: (2021)
by: Cinà, Antonio Emanuele, et al.
Published: (2021)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
by: Yao, Duanyi, et al.
Published: (2026)
by: Yao, Duanyi, et al.
Published: (2026)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
by: Vyas, Sanyam, et al.
Published: (2024)
by: Vyas, Sanyam, et al.
Published: (2024)
Fusing Pruned and Backdoored Models: Optimal Transport-based Data-free Backdoor Mitigation
by: Lin, Weilin, et al.
Published: (2024)
by: Lin, Weilin, et al.
Published: (2024)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
by: Yamashita, Tomoya, et al.
Published: (2025)
by: Yamashita, Tomoya, et al.
Published: (2025)
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Improving Intrusion Detection with Domain-Invariant Representation Learning in Latent Space
by: Roy, Padmaksha, et al.
Published: (2023)
by: Roy, Padmaksha, et al.
Published: (2023)
Byzantine-Robust Federated Learning: An Overview With Focus on Developing Sybil-based Attacks to Backdoor Augmented Secure Aggregation Protocols
by: Deshmukh, Atharv
Published: (2024)
by: Deshmukh, Atharv
Published: (2024)
Seal Your Backdoor with Variational Defense
by: Sabolić, Ivan, et al.
Published: (2025)
by: Sabolić, Ivan, et al.
Published: (2025)
Persistent Backdoor Attacks in Continual Learning
by: Guo, Zhen, et al.
Published: (2024)
by: Guo, Zhen, et al.
Published: (2024)
Backdooring Masked Diffusion Language Models
by: Cao, Daniel Yiming, et al.
Published: (2026)
by: Cao, Daniel Yiming, et al.
Published: (2026)
Similar Items
-
Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?
by: Ito, Akira, et al.
Published: (2025) -
A Channel-Triggered Backdoor Attack on Wireless Semantic Image Reconstruction
by: Wan, Jialin, et al.
Published: (2025) -
Hardware-Triggered Backdoors
by: Möller, Jonas, et al.
Published: (2026) -
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026) -
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
by: Ashcraft, Chace, et al.
Published: (2025)