Distillation Robustifies Unlearning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Bruce W., Foote, Addie, Infanger, Alex, Shor, Leni, Kamath, Harish, Goldman-Wetzler, Jacob, Woodworth, Bryce, Cloud, Alex, Turner, Alexander Matt |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025)
by: Drori, Jacob, et al.
Published: (2025)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025)
by: Azarbal, Ariana, et al.
Published: (2025)
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026)
by: Marklund, Henrik, et al.
Published: (2026)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
The Persian Rug: solving toy models of superposition using large-scale symmetries
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Active propulsion noise shaping for multi-rotor aircraft localization
by: Serussi, Gabriele, et al.
Published: (2024)
by: Serussi, Gabriele, et al.
Published: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
by: Irpan, Alex, et al.
Published: (2025)
by: Irpan, Alex, et al.
Published: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training
by: Goldman, Richard, et al.
Published: (2025)
by: Goldman, Richard, et al.
Published: (2025)
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
by: Oesterling, Alex, et al.
Published: (2023)
by: Oesterling, Alex, et al.
Published: (2023)
Verifying Machine Unlearning with Explainable AI
by: Vidal, Àlex Pujol, et al.
Published: (2024)
by: Vidal, Àlex Pujol, et al.
Published: (2024)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
Robustifying and Selecting Cohort-Appropriate Prognostic Models under Distributional Shifts
by: Bertsimas, Dimitris, et al.
Published: (2026)
by: Bertsimas, Dimitris, et al.
Published: (2026)
HyperCLOVA X THINK Technical Report
by: NAVER Cloud HyperCLOVA X Team
Published: (2025)
by: NAVER Cloud HyperCLOVA X Team
Published: (2025)
Machine Unlearning under Overparameterization
by: Block, Jacob L., et al.
Published: (2025)
by: Block, Jacob L., et al.
Published: (2025)
ADAPT to Robustify Prompt Tuning Vision Transformers
by: Eskandar, Masih, et al.
Published: (2024)
by: Eskandar, Masih, et al.
Published: (2024)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
QuickDrop: Efficient Federated Unlearning by Integrated Dataset Distillation
by: Dhasade, Akash, et al.
Published: (2023)
by: Dhasade, Akash, et al.
Published: (2023)
ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs
by: Maheri, Mohammad M, et al.
Published: (2025)
by: Maheri, Mohammad M, et al.
Published: (2025)
Robustifying Safety-Aligned Large Language Models through Clean Data Curation
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
LLM Unlearning via Neural Activation Redirection
by: Shen, William F., et al.
Published: (2025)
by: Shen, William F., et al.
Published: (2025)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
by: Shilov, Igor, et al.
Published: (2025)
by: Shilov, Igor, et al.
Published: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Machine Unlearning Fails to Remove Data Poisoning Attacks
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
by: Hu, Zizhao, et al.
Published: (2026)
by: Hu, Zizhao, et al.
Published: (2026)
UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Unsupervised Elicitation of Language Models
by: Wen, Jiaxin, et al.
Published: (2025)
by: Wen, Jiaxin, et al.
Published: (2025)
Relation-based Counterfactual Data Augmentation and Contrastive Learning for Robustifying Natural Language Inference Models
by: Yang, Heerin, et al.
Published: (2024)
by: Yang, Heerin, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
HyperCLOVA X 8B Omni
by: NAVER Cloud HyperCLOVA X Team
Published: (2026)
by: NAVER Cloud HyperCLOVA X Team
Published: (2026)
Inteligencia Artificial jurídica y el desafío de la veracidad: análisis de alucinaciones, optimización de RAG y principios para una integración responsable
by: Dantart, Alex
Published: (2025)
by: Dantart, Alex
Published: (2025)
Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability
by: Grzankowski, Alex
Published: (2024)
by: Grzankowski, Alex
Published: (2024)
Reinforcement Learning via Auxiliary Task Distillation
by: Harish, Abhinav Narayan, et al.
Published: (2024)
by: Harish, Abhinav Narayan, et al.
Published: (2024)
Distill, Forget, Repeat: A Framework for Continual Unlearning in Text-to-Image Diffusion Models
by: George, Naveen, et al.
Published: (2025)
by: George, Naveen, et al.
Published: (2025)
Relationship-Aware Safety Unlearning for Multimodal LLMs
by: Anilkumar, Vishnu Narayanan, et al.
Published: (2026)
by: Anilkumar, Vishnu Narayanan, et al.
Published: (2026)
Similar Items
-
Output Supervision Can Obfuscate the Chain of Thought
by: Drori, Jacob, et al.
Published: (2025) -
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024) -
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025) -
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026) -
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)