GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Yue, Wang, Qizhou, Liu, Feng, Huang, Wei, Du, Yali, Du, Xiaojiang, Han, Bo |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM Unlearning with LLM Beliefs
par: Li, Kemou, et autres
Publié: (2025)
par: Li, Kemou, et autres
Publié: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
par: Xu, Xiaoyu, et autres
Publié: (2025)
par: Xu, Xiaoyu, et autres
Publié: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
par: Nawrot, Piotr, et autres
Publié: (2025)
par: Nawrot, Piotr, et autres
Publié: (2025)
Explainable LLM Unlearning Through Reasoning
par: Liao, Junfeng, et autres
Publié: (2026)
par: Liao, Junfeng, et autres
Publié: (2026)
Safe Reinforcement Learning with Free-form Natural Language Constraints and Pre-Trained Language Models
par: Lou, Xingzhou, et autres
Publié: (2024)
par: Lou, Xingzhou, et autres
Publié: (2024)
Conformal Linguistic Calibration: Trading-off between Factuality and Specificity
par: Jiang, Zhengping, et autres
Publié: (2025)
par: Jiang, Zhengping, et autres
Publié: (2025)
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
par: Zhao, Yang, et autres
Publié: (2026)
par: Zhao, Yang, et autres
Publié: (2026)
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
par: Zhang, Zhicheng, et autres
Publié: (2025)
par: Zhang, Zhicheng, et autres
Publié: (2025)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
par: Du, Tianqi, et autres
Publié: (2025)
par: Du, Tianqi, et autres
Publié: (2025)
TOFU: A Task of Fictitious Unlearning for LLMs
par: Maini, Pratyush, et autres
Publié: (2024)
par: Maini, Pratyush, et autres
Publié: (2024)
Leverage Unlearning to Sanitize LLMs
par: Boutet, Antoine, et autres
Publié: (2025)
par: Boutet, Antoine, et autres
Publié: (2025)
MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
par: Wang, Ke, et autres
Publié: (2025)
par: Wang, Ke, et autres
Publié: (2025)
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
par: Liu, Wei, et autres
Publié: (2025)
par: Liu, Wei, et autres
Publié: (2025)
Textual Unlearning Gives a False Sense of Unlearning
par: Du, Jiacheng, et autres
Publié: (2024)
par: Du, Jiacheng, et autres
Publié: (2024)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
par: Qin, Haotong, et autres
Publié: (2024)
par: Qin, Haotong, et autres
Publié: (2024)
All Language Models Large and Small
par: Chen, Zhixun, et autres
Publié: (2024)
par: Chen, Zhixun, et autres
Publié: (2024)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
par: Liu, Wei, et autres
Publié: (2026)
par: Liu, Wei, et autres
Publié: (2026)
Machine Unlearning of Pre-trained Large Language Models
par: Yao, Jin, et autres
Publié: (2024)
par: Yao, Jin, et autres
Publié: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
par: Kadhe, Swanand Ravindra, et autres
Publié: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
par: Wu, Ruihan, et autres
Publié: (2025)
par: Wu, Ruihan, et autres
Publié: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
par: Farashah, Alireza Dehghanpour, et autres
Publié: (2026)
par: Farashah, Alireza Dehghanpour, et autres
Publié: (2026)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
par: Lu, Taiming, et autres
Publié: (2024)
par: Lu, Taiming, et autres
Publié: (2024)
Bridging the Gap Between Preference Alignment and Machine Unlearning
par: Feng, Xiaohua, et autres
Publié: (2025)
par: Feng, Xiaohua, et autres
Publié: (2025)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
par: Niu, Peizhi, et autres
Publié: (2025)
par: Niu, Peizhi, et autres
Publié: (2025)
Exploring Accuracy-Fairness Trade-off in Large Language Models
par: Zhang, Qingquan, et autres
Publié: (2024)
par: Zhang, Qingquan, et autres
Publié: (2024)
Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
par: Liu, Wei, et autres
Publié: (2026)
par: Liu, Wei, et autres
Publié: (2026)
Safe Multi-agent Reinforcement Learning with Natural Language Constraints
par: Wang, Ziyan, et autres
Publié: (2024)
par: Wang, Ziyan, et autres
Publié: (2024)
CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment
par: Guo, Siyuan, et autres
Publié: (2026)
par: Guo, Siyuan, et autres
Publié: (2026)
Towards Transfer Unlearning: Empirical Evidence of Cross-Domain Bias Mitigation
par: Lu, Huimin, et autres
Publié: (2024)
par: Lu, Huimin, et autres
Publié: (2024)
PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
par: Chen, Xiaoyi, et autres
Publié: (2026)
par: Chen, Xiaoyi, et autres
Publié: (2026)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
par: Cha, Sungmin, et autres
Publié: (2024)
par: Cha, Sungmin, et autres
Publié: (2024)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
par: Zhang, Xuan, et autres
Publié: (2024)
par: Zhang, Xuan, et autres
Publié: (2024)
How well can off-the-shelf LLMs elucidate molecular structures from mass spectra using chain-of-thought reasoning?
par: Wang, Yufeng, et autres
Publié: (2026)
par: Wang, Yufeng, et autres
Publié: (2026)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
par: Fan, Chongyu, et autres
Publié: (2025)
par: Fan, Chongyu, et autres
Publié: (2025)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
par: Wei, Rongzhe, et autres
Publié: (2025)
par: Wei, Rongzhe, et autres
Publié: (2025)
A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
par: Feng, Xiaohua, et autres
Publié: (2025)
par: Feng, Xiaohua, et autres
Publié: (2025)
LANCET: Neural Intervention via Structural Entropy for Mitigating Faithfulness Hallucinations in LLMs
par: Wang, Chenxu, et autres
Publié: (2026)
par: Wang, Chenxu, et autres
Publié: (2026)
A General Framework to Enhance Fine-tuning-based LLM Unlearning
par: Ren, Jie, et autres
Publié: (2025)
par: Ren, Jie, et autres
Publié: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
par: Jaiswal, Ajay, et autres
Publié: (2023)
par: Jaiswal, Ajay, et autres
Publié: (2023)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
par: Liu, Bing, et autres
Publié: (2026)
par: Liu, Bing, et autres
Publié: (2026)
Documents similaires
-
LLM Unlearning with LLM Beliefs
par: Li, Kemou, et autres
Publié: (2025) -
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
par: Xu, Xiaoyu, et autres
Publié: (2025) -
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
par: Nawrot, Piotr, et autres
Publié: (2025) -
Explainable LLM Unlearning Through Reasoning
par: Liao, Junfeng, et autres
Publié: (2026) -
Safe Reinforcement Learning with Free-form Natural Language Constraints and Pre-Trained Language Models
par: Lou, Xingzhou, et autres
Publié: (2024)