Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Aravindan, Ashwath Vaithinathan, Jha, Abha, Salaway, Matthew, Bhide, Atharva Sandeep, Yaldiz, Duygu Nur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
by: Jha, Abha, et al.
Published: (2025)
by: Jha, Abha, et al.
Published: (2025)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026)
Code-Driven Planning in Grid Worlds with Large Language Models
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
by: Jha, Abha, et al.
Published: (2026)
by: Jha, Abha, et al.
Published: (2026)
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
by: Wang, Shanmin, et al.
Published: (2025)
by: Wang, Shanmin, et al.
Published: (2025)
CroMo-Mixup: Augmenting Cross-Model Representations for Continual Self-Supervised Learning
by: Mushtaq, Erum, et al.
Published: (2024)
by: Mushtaq, Erum, et al.
Published: (2024)
Trigger without Trace: Towards Stealthy Backdoor Attack on Text-to-Image Diffusion Models
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026)
by: Abad, Gorka, et al.
Published: (2026)
Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Robust MLLM Unlearning via Visual Knowledge Distillation
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Backdoor Unlearning by Linear Task Decomposition
by: Abdelraheem, Amel, et al.
Published: (2025)
by: Abdelraheem, Amel, et al.
Published: (2025)
Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling
by: Li, Zida, et al.
Published: (2026)
by: Li, Zida, et al.
Published: (2026)
Adversarially Guided Stateful Defense Against Backdoor Attacks in Federated Deep Learning
by: Ali, Hassan, et al.
Published: (2024)
by: Ali, Hassan, et al.
Published: (2024)
Probing Unlearned Diffusion Models: A Transferable Adversarial Attack Perspective
by: Han, Xiaoxuan, et al.
Published: (2024)
by: Han, Xiaoxuan, et al.
Published: (2024)
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
by: Mekonnen, Kidist Amde, et al.
Published: (2024)
TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models
by: Xiang, Qianlong, et al.
Published: (2026)
by: Xiang, Qianlong, et al.
Published: (2026)
Adversarial Backdoor Defense in CLIP
by: Kuang, Junhao, et al.
Published: (2024)
by: Kuang, Junhao, et al.
Published: (2024)
Using Interleaved Ensemble Unlearning to Keep Backdoors at Bay for Finetuning Vision Transformers
by: Li, Zeyu Michael
Published: (2024)
by: Li, Zeyu Michael
Published: (2024)
Dynamic Guidance Adversarial Distillation with Enhanced Teacher Knowledge
by: Park, Hyejin, et al.
Published: (2024)
by: Park, Hyejin, et al.
Published: (2024)
Teacher-Guided Student Self-Knowledge Distillation Using Diffusion Model
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Is It Possible to Backdoor Face Forgery Detection with Natural Triggers?
by: Han, Xiaoxuan, et al.
Published: (2023)
by: Han, Xiaoxuan, et al.
Published: (2023)
Explainability-Driven Leaf Disease Classification Using Adversarial Training and Knowledge Distillation
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
Adversarial Concept Distillation for One-Step Diffusion Personalization
by: Yang, Yixiong, et al.
Published: (2025)
by: Yang, Yixiong, et al.
Published: (2025)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2024)
by: Wang, Zhongqi, et al.
Published: (2024)
Unlearning Concepts from Text-to-Video Diffusion Models
by: Liu, Shiqi, et al.
Published: (2024)
by: Liu, Shiqi, et al.
Published: (2024)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
by: Zhang, Yimeng, et al.
Published: (2024)
by: Zhang, Yimeng, et al.
Published: (2024)
Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning
by: Moon, Saemi, et al.
Published: (2024)
by: Moon, Saemi, et al.
Published: (2024)
Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor Activeness
by: Lin, Weilin, et al.
Published: (2024)
by: Lin, Weilin, et al.
Published: (2024)
Data-free Knowledge Distillation with Diffusion Models
by: Qi, Xiaohua, et al.
Published: (2025)
by: Qi, Xiaohua, et al.
Published: (2025)
Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
by: Koma, Arian Komaei, et al.
Published: (2026)
by: Koma, Arian Komaei, et al.
Published: (2026)
Backdoor Attack with Sparse and Invisible Trigger
by: Gao, Yinghua, et al.
Published: (2023)
by: Gao, Yinghua, et al.
Published: (2023)
Towards Adversarial Robustness And Backdoor Mitigation in SSL
by: Satpathy, Aryan, et al.
Published: (2024)
by: Satpathy, Aryan, et al.
Published: (2024)
Seal2Real: Prompt Prior Learning on Diffusion Model for Unsupervised Document Seal Data Generation and Realisation
by: Yan, Mingfu, et al.
Published: (2023)
by: Yan, Mingfu, et al.
Published: (2023)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Twin Trigger Generative Networks for Backdoor Attacks against Object Detection
by: Li, Zhiying, et al.
Published: (2024)
by: Li, Zhiying, et al.
Published: (2024)
Invisible Backdoor Triggers in Image Editing Model via Deep Watermarking
by: Chen, Yu-Feng, et al.
Published: (2025)
by: Chen, Yu-Feng, et al.
Published: (2025)
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
by: Teng, Wenbin, et al.
Published: (2025)
by: Teng, Wenbin, et al.
Published: (2025)
Similar Items
-
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
by: Jha, Abha, et al.
Published: (2025) -
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025) -
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026) -
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2026) -
Code-Driven Planning in Grid Worlds with Large Language Models
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)