How to Protect Models against Adversarial Unlearning?
Fuente:
arXiv
Saved in:
| Main Authors: | Jasiorski, Patryk, Klonowski, Marek, Woźniak, Michał |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlearning-based sliding window for continual learning under concept drift
by: Wozniak, Michal, et al.
Published: (2026)
by: Wozniak, Michal, et al.
Published: (2026)
ROKA: Robust Knowledge Unlearning against Adversaries
by: Shin, Jinmyeong, et al.
Published: (2026)
by: Shin, Jinmyeong, et al.
Published: (2026)
Discriminative Adversarial Unlearning
by: Sharma, Rohan, et al.
Published: (2024)
by: Sharma, Rohan, et al.
Published: (2024)
EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions
by: Bartkowiak, Patryk, et al.
Published: (2025)
by: Bartkowiak, Patryk, et al.
Published: (2025)
Zero-Shot Machine Unlearning with Proxy Adversarial Data Generation
by: Chen, Huiqiang, et al.
Published: (2025)
by: Chen, Huiqiang, et al.
Published: (2025)
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
by: Zhao, Mengnan, et al.
Published: (2024)
by: Zhao, Mengnan, et al.
Published: (2024)
A Natural Gas Consumption Forecasting System for Continual Learning Scenarios based on Hoeffding Trees with Change Point Detection Mechanism
by: Svoboda, Radek, et al.
Published: (2023)
by: Svoboda, Radek, et al.
Published: (2023)
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
by: Walkowiak, Paweł, et al.
Published: (2025)
by: Walkowiak, Paweł, et al.
Published: (2025)
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
by: Spiegel, Michal, et al.
Published: (2025)
by: Spiegel, Michal, et al.
Published: (2025)
How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning
by: Chen, Jiangwei, et al.
Published: (2026)
by: Chen, Jiangwei, et al.
Published: (2026)
Towards consistency of rule-based explainer and black box model -- fusion of rule induction and XAI-based feature importance
by: Kozielski, Michał, et al.
Published: (2024)
by: Kozielski, Michał, et al.
Published: (2024)
HyConEx: Hypernetwork classifier with counterfactual explanations for tabular data
by: Marszałek, Patryk, et al.
Published: (2025)
by: Marszałek, Patryk, et al.
Published: (2025)
How Do Diffusion Models Improve Adversarial Robustness?
by: Yuezhang, Liu, et al.
Published: (2025)
by: Yuezhang, Liu, et al.
Published: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
by: Winninger, Thomas, et al.
Published: (2025)
by: Winninger, Thomas, et al.
Published: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
by: Rossolini, Giulio
Published: (2026)
by: Rossolini, Giulio
Published: (2026)
Adapt then Unlearn: Exploring Parameter Space Semantics for Unlearning in Generative Adversarial Networks
by: Tiwary, Piyush, et al.
Published: (2023)
by: Tiwary, Piyush, et al.
Published: (2023)
Ready2Unlearn: A Learning-Time Approach for Preparing Models with Future Unlearning Readiness
by: Duan, Hanyu, et al.
Published: (2025)
by: Duan, Hanyu, et al.
Published: (2025)
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
by: Yamabe, Shojiro, et al.
Published: (2024)
by: Yamabe, Shojiro, et al.
Published: (2024)
Deep Unlearn: Benchmarking Machine Unlearning for Image Classification
by: Cadet, Xavier F., et al.
Published: (2024)
by: Cadet, Xavier F., et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Fast Adversarial Training against Sparse Attacks Requires Loss Smoothing
by: Zhong, Xuyang, et al.
Published: (2025)
by: Zhong, Xuyang, et al.
Published: (2025)
Certified Robustness against Sparse Adversarial Perturbations via Data Localization
by: Pal, Ambar, et al.
Published: (2024)
by: Pal, Ambar, et al.
Published: (2024)
Holistic Continual Learning under Concept Drift with Adaptive Memory Realignment
by: Ashrafee, Alif, et al.
Published: (2025)
by: Ashrafee, Alif, et al.
Published: (2025)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
Unlearning Information Bottleneck: Machine Unlearning of Systematic Patterns and Biases
by: Han, Ling, et al.
Published: (2024)
by: Han, Ling, et al.
Published: (2024)
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
by: Grant, Satchel, et al.
Published: (2026)
by: Grant, Satchel, et al.
Published: (2026)
Deep Dictionary-Free Method for Identifying Linear Model of Nonlinear System with Input Delay
by: Valábek, Patrik, et al.
Published: (2025)
by: Valábek, Patrik, et al.
Published: (2025)
Contract And Conquer: How to Provably Compute Adversarial Examples for a Black-Box Model?
by: Chistyakova, Anna, et al.
Published: (2026)
by: Chistyakova, Anna, et al.
Published: (2026)
Auditing Approximate Machine Unlearning for Differentially Private Models
by: Gu, Yuechun, et al.
Published: (2025)
by: Gu, Yuechun, et al.
Published: (2025)
Auditing Language Model Unlearning via Information Decomposition
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
Federated Knowledge Graph Unlearning via Diffusion Model
by: Liu, Bingchen, et al.
Published: (2024)
by: Liu, Bingchen, et al.
Published: (2024)
Distillation Robustifies Unlearning
by: Lee, Bruce W., et al.
Published: (2025)
by: Lee, Bruce W., et al.
Published: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
by: Lin, Yujie, et al.
Published: (2026)
by: Lin, Yujie, et al.
Published: (2026)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
by: Tang, Haoyu, et al.
Published: (2024)
by: Tang, Haoyu, et al.
Published: (2024)
XSub: Explanation-Driven Adversarial Attack against Blackbox Classifiers via Feature Substitution
by: Vu, Kiana, et al.
Published: (2024)
by: Vu, Kiana, et al.
Published: (2024)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
by: Kawakami, Tatsuki, et al.
Published: (2025)
by: Kawakami, Tatsuki, et al.
Published: (2025)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Similar Items
-
Unlearning-based sliding window for continual learning under concept drift
by: Wozniak, Michal, et al.
Published: (2026) -
ROKA: Robust Knowledge Unlearning against Adversaries
by: Shin, Jinmyeong, et al.
Published: (2026) -
Discriminative Adversarial Unlearning
by: Sharma, Rohan, et al.
Published: (2024) -
EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions
by: Bartkowiak, Patryk, et al.
Published: (2025) -
Zero-Shot Machine Unlearning with Proxy Adversarial Data Generation
by: Chen, Huiqiang, et al.
Published: (2025)