Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cha, Sungmin, Cho, Sungjun, Hwang, Dasol, Lee, Moontae |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
by: Cha, Sungmin, et al.
Published: (2023)
by: Cha, Sungmin, et al.
Published: (2023)
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
by: Cho, Sungjun, et al.
Published: (2025)
by: Cho, Sungjun, et al.
Published: (2025)
Knowledge Vector Weakening: Efficient Training-free Unlearning for Large Vision-Language Models
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
by: Han, Sungjun, et al.
Published: (2024)
by: Han, Sungjun, et al.
Published: (2024)
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
by: Cha, Sungmin, et al.
Published: (2025)
by: Cha, Sungmin, et al.
Published: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Hyperparameters in Continual Learning: A Reality Check
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
by: Wei, Rongzhe, et al.
Published: (2025)
by: Wei, Rongzhe, et al.
Published: (2025)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
Leverage Unlearning to Sanitize LLMs
by: Boutet, Antoine, et al.
Published: (2025)
by: Boutet, Antoine, et al.
Published: (2025)
Are We Truly Forgetting? A Critical Re-examination of Machine Unlearning Evaluation Protocols
by: Kim, Yongwoo, et al.
Published: (2025)
by: Kim, Yongwoo, et al.
Published: (2025)
CURaTE: Continual Unlearning in Real Time with Ensured Preservation of LLM Knowledge
by: Bae, Seyun, et al.
Published: (2026)
by: Bae, Seyun, et al.
Published: (2026)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
by: Cha, Sungmin, et al.
Published: (2022)
by: Cha, Sungmin, et al.
Published: (2022)
TOFU: A Task of Fictitious Unlearning for LLMs
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
by: Lim, Sungjun, et al.
Published: (2026)
by: Lim, Sungjun, et al.
Published: (2026)
Why Alignment Must Precede Distillation: A Minimal Working Explanation
by: Cha, Sungmin, et al.
Published: (2025)
by: Cha, Sungmin, et al.
Published: (2025)
Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
by: Han, Sungmin, et al.
Published: (2025)
by: Han, Sungmin, et al.
Published: (2025)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
Efficient Knowledge Injection in LLMs via Self-Distillation
by: Kujanpää, Kalle, et al.
Published: (2024)
by: Kujanpää, Kalle, et al.
Published: (2024)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
by: Lin, Yujie, et al.
Published: (2026)
by: Lin, Yujie, et al.
Published: (2026)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models
by: Mekala, Anmol, et al.
Published: (2024)
by: Mekala, Anmol, et al.
Published: (2024)
GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping
by: Dahal, Chahana, et al.
Published: (2026)
by: Dahal, Chahana, et al.
Published: (2026)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
by: Yamashita, Tomoya, et al.
Published: (2025)
by: Yamashita, Tomoya, et al.
Published: (2025)
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
by: Yeats, Eric, et al.
Published: (2025)
by: Yeats, Eric, et al.
Published: (2025)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
by: Wu, Qinyuan, et al.
Published: (2024)
by: Wu, Qinyuan, et al.
Published: (2024)
Towards Realistic Incremental Scenario in Class Incremental Semantic Segmentation
by: Kwak, Jihwan, et al.
Published: (2024)
by: Kwak, Jihwan, et al.
Published: (2024)
Similar Items
-
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
by: Cha, Sungmin, et al.
Published: (2023) -
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
by: Cho, Sungjun, et al.
Published: (2025) -
Knowledge Vector Weakening: Efficient Training-free Unlearning for Large Vision-Language Models
by: Kim, Yejin, et al.
Published: (2026) -
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024) -
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)