Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Tingxu, Sun, Weisong, Ding, Ziqi, Fang, Chunrong, Qian, Hanwei, Li, Jiaxun, Chen, Zhenyu, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Distillation Mitigate Backdoor Attacks in Pre-trained Encoders?
by: Han, TIngxu, et al.
Published: (2024)
by: Han, TIngxu, et al.
Published: (2024)
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024)
by: Chen, Yuchen, et al.
Published: (2024)
Eliminating Backdoors in Neural Code Models for Secure Code Understanding
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
by: Zhao, Tianhang, et al.
Published: (2025)
by: Zhao, Tianhang, et al.
Published: (2025)
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
by: Peng, Zuquan, et al.
Published: (2025)
by: Peng, Zuquan, et al.
Published: (2025)
ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders
by: Jiang, Yongqi, et al.
Published: (2026)
by: Jiang, Yongqi, et al.
Published: (2026)
SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-supervised Learning
by: Lv, Peizhuo, et al.
Published: (2022)
by: Lv, Peizhuo, et al.
Published: (2022)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
Probe-Me-Not: Protecting Pre-trained Encoders from Malicious Probing
by: Ding, Ruyi, et al.
Published: (2024)
by: Ding, Ruyi, et al.
Published: (2024)
Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models
by: Chen, Linzhi, et al.
Published: (2025)
by: Chen, Linzhi, et al.
Published: (2025)
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
by: Chen, Shenyang, et al.
Published: (2026)
by: Chen, Shenyang, et al.
Published: (2026)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
by: Kong, Cong, et al.
Published: (2024)
by: Kong, Cong, et al.
Published: (2024)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
by: Ding, Zikang, et al.
Published: (2026)
by: Ding, Zikang, et al.
Published: (2026)
DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
by: Chen, Yuchen, et al.
Published: (2026)
by: Chen, Yuchen, et al.
Published: (2026)
Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets
by: Zhang, Yechao, et al.
Published: (2025)
by: Zhang, Yechao, et al.
Published: (2025)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
Concept-Guided Backdoor Attack on Vision Language Models
by: Shen, Haoyu, et al.
Published: (2025)
by: Shen, Haoyu, et al.
Published: (2025)
Pre-trained Encoder Inference: Revealing Upstream Encoders In Downstream Machine Learning Services
by: Fu, Shaopeng, et al.
Published: (2024)
by: Fu, Shaopeng, et al.
Published: (2024)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code Abstraction
by: Xiao, Yuan, et al.
Published: (2025)
by: Xiao, Yuan, et al.
Published: (2025)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
On the Weaknesses of Backdoor-based Model Watermarking: An Information-theoretic Perspective
by: Hu, Aoting, et al.
Published: (2024)
by: Hu, Aoting, et al.
Published: (2024)
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
by: Shen, Guangyu, et al.
Published: (2025)
by: Shen, Guangyu, et al.
Published: (2025)
Membership Inference for Contrastive Pre-training Models with Text-only PII Queries
by: Cheng, Ruoxi, et al.
Published: (2026)
by: Cheng, Ruoxi, et al.
Published: (2026)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
A Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories
by: Ding, Ziqi, et al.
Published: (2025)
by: Ding, Ziqi, et al.
Published: (2025)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
MARS: A Malignity-Aware Backdoor Defense in Federated Learning
by: Wan, Wei, et al.
Published: (2025)
by: Wan, Wei, et al.
Published: (2025)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Lightweight and Fast Backdoor Model Detection
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
Demonstration Attack against In-Context Learning for Code Intelligence
by: Ge, Yifei, et al.
Published: (2024)
by: Ge, Yifei, et al.
Published: (2024)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
TokenMark: A Modality-Agnostic Watermark for Pre-trained Transformers
by: Xu, Hengyuan, et al.
Published: (2024)
by: Xu, Hengyuan, et al.
Published: (2024)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
PBP: Post-training Backdoor Purification for Malware Classifiers
by: Nguyen, Dung Thuy, et al.
Published: (2024)
by: Nguyen, Dung Thuy, et al.
Published: (2024)
SFIBA: Spatial-based Full-target Invisible Backdoor Attacks
by: Yin, Yangxu, et al.
Published: (2025)
by: Yin, Yangxu, et al.
Published: (2025)
Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation Zone
by: Xiao, Yuan, et al.
Published: (2022)
by: Xiao, Yuan, et al.
Published: (2022)
Similar Items
-
Can Distillation Mitigate Backdoor Attacks in Pre-trained Encoders?
by: Han, TIngxu, et al.
Published: (2024) -
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024) -
Eliminating Backdoors in Neural Code Models for Secure Code Understanding
by: Sun, Weisong, et al.
Published: (2024) -
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
by: Zhao, Tianhang, et al.
Published: (2025) -
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
by: Zhao, Xiangyu, et al.
Published: (2024)