Rethinking Machine Unlearning for Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Sijia, Yao, Yuanshun, Jia, Jinghan, Casper, Stephen, Baracaldo, Nathalie, Hase, Peter, Yao, Yuguang, Liu, Chris Yuhao, Xu, Xiaojun, Li, Hang, Varshney, Kush R., Bansal, Mohit, Koyejo, Sanmi, Liu, Yang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability
par: Jia, Jinghan, et autres
Publié: (2025)
par: Jia, Jinghan, et autres
Publié: (2025)
Label Smoothing Improves Machine Unlearning
par: Di, Zonglin, et autres
Publié: (2024)
par: Di, Zonglin, et autres
Publié: (2024)
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
par: Wang, Changsheng, et autres
Publié: (2025)
par: Wang, Changsheng, et autres
Publié: (2025)
Large Language Model Unlearning
par: Yao, Yuanshun, et autres
Publié: (2023)
par: Yao, Yuanshun, et autres
Publié: (2023)
Model Sparsity Can Simplify Machine Unlearning
par: Jia, Jinghan, et autres
Publié: (2023)
par: Jia, Jinghan, et autres
Publié: (2023)
WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models
par: Jia, Jinghan, et autres
Publié: (2024)
par: Jia, Jinghan, et autres
Publié: (2024)
Robust Multi-bit Text Watermark with LLM-based Paraphrasers
par: Xu, Xiaojun, et autres
Publié: (2024)
par: Xu, Xiaojun, et autres
Publié: (2024)
UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
par: Zhang, Yihua, et autres
Publié: (2024)
par: Zhang, Yihua, et autres
Publié: (2024)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
par: Wang, Changsheng, et autres
Publié: (2025)
par: Wang, Changsheng, et autres
Publié: (2025)
Distributional Machine Unlearning via Selective Data Removal
par: Allouah, Youssef, et autres
Publié: (2025)
par: Allouah, Youssef, et autres
Publié: (2025)
Learning to Watermark LLM-generated Text via Reinforcement Learning
par: Xu, Xiaojun, et autres
Publié: (2024)
par: Xu, Xiaojun, et autres
Publié: (2024)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
par: Fan, Chongyu, et autres
Publié: (2024)
par: Fan, Chongyu, et autres
Publié: (2024)
EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
par: Jia, Jinghan, et autres
Publié: (2025)
par: Jia, Jinghan, et autres
Publié: (2025)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
par: Allouah, Youssef, et autres
Publié: (2024)
par: Allouah, Youssef, et autres
Publié: (2024)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
par: Chen, Yiwei, et autres
Publié: (2025)
par: Chen, Yiwei, et autres
Publié: (2025)
On the Cause of Unfairness: A Training Sample Perspective
par: Yao, Yuanshun, et autres
Publié: (2023)
par: Yao, Yuanshun, et autres
Publié: (2023)
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
par: Sun, Changchang, et autres
Publié: (2024)
par: Sun, Changchang, et autres
Publié: (2024)
Adversarial Watermarking for Face Recognition
par: Yao, Yuguang, et autres
Publié: (2024)
par: Yao, Yuguang, et autres
Publié: (2024)
Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
par: Sun, Yuhao, et autres
Publié: (2025)
par: Sun, Yuhao, et autres
Publié: (2025)
Certified Unlearning for Neural Networks
par: Koloskova, Anastasia, et autres
Publié: (2025)
par: Koloskova, Anastasia, et autres
Publié: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
par: Zhuang, Haomin, et autres
Publié: (2024)
par: Zhuang, Haomin, et autres
Publié: (2024)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
par: Lang, Yicheng, et autres
Publié: (2025)
par: Lang, Yicheng, et autres
Publié: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
par: Patil, Vaidehi, et autres
Publié: (2025)
par: Patil, Vaidehi, et autres
Publié: (2025)
Hide and Seek: How Does Watermarking Impact Face Recognition?
par: Yao, Yuguang, et autres
Publié: (2024)
par: Yao, Yuguang, et autres
Publié: (2024)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
par: Fan, Chongyu, et autres
Publié: (2025)
par: Fan, Chongyu, et autres
Publié: (2025)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
par: Jia, Jinghan, et autres
Publié: (2024)
par: Jia, Jinghan, et autres
Publié: (2024)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
par: Wang, Changsheng, et autres
Publié: (2025)
par: Wang, Changsheng, et autres
Publié: (2025)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
par: Stengel-Eskin, Elias, et autres
Publié: (2024)
par: Stengel-Eskin, Elias, et autres
Publié: (2024)
Teaching Models to Balance Resisting and Accepting Persuasion
par: Stengel-Eskin, Elias, et autres
Publié: (2024)
par: Stengel-Eskin, Elias, et autres
Publié: (2024)
Causally Inspired Regularization Enables Domain General Representations
par: Salaudeen, Olawale, et autres
Publié: (2024)
par: Salaudeen, Olawale, et autres
Publié: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
par: Robertson, Zachary, et autres
Publié: (2025)
par: Robertson, Zachary, et autres
Publié: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
par: Vo, Truong, et autres
Publié: (2025)
par: Vo, Truong, et autres
Publié: (2025)
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
par: Shah, Raj Sanjay, et autres
Publié: (2026)
par: Shah, Raj Sanjay, et autres
Publié: (2026)
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
par: Zhang, Yimeng, et autres
Publié: (2023)
par: Zhang, Yimeng, et autres
Publié: (2023)
Decolonial AI Alignment: Openness, Viśe\d{s}a-Dharma, and Including Excluded Knowledges
par: Varshney, Kush R.
Publié: (2023)
par: Varshney, Kush R.
Publié: (2023)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
par: Varshney, Kush R.
Publié: (2025)
par: Varshney, Kush R.
Publié: (2025)
An Algebraic Exposition of the Theory of Dyadic Morality
par: Varshney, Kush R.
Publié: (2026)
par: Varshney, Kush R.
Publié: (2026)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
par: Fan, Chongyu, et autres
Publié: (2024)
par: Fan, Chongyu, et autres
Publié: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
par: Zhang, Yimeng, et autres
Publié: (2024)
par: Zhang, Yimeng, et autres
Publié: (2024)
Documents similaires
-
Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability
par: Jia, Jinghan, et autres
Publié: (2025) -
Label Smoothing Improves Machine Unlearning
par: Di, Zonglin, et autres
Publié: (2024) -
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
par: Wang, Changsheng, et autres
Publié: (2025) -
Large Language Model Unlearning
par: Yao, Yuanshun, et autres
Publié: (2023) -
Model Sparsity Can Simplify Machine Unlearning
par: Jia, Jinghan, et autres
Publié: (2023)