Offset Unlearning for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, James Y., Zhou, Wenxuan, Wang, Fei, Morstatter, Fred, Zhang, Sheng, Poon, Hoifung, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
von: Askari, Hadi, et al.
Veröffentlicht: (2025)
von: Askari, Hadi, et al.
Veröffentlicht: (2025)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
Contrastive Instruction Tuning
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2024)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2024)
Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongxiang, et al.
Veröffentlicht: (2025)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
Instructional Fingerprinting of Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
von: Salinas, Abel, et al.
Veröffentlicht: (2024)
von: Salinas, Abel, et al.
Veröffentlicht: (2024)
Large Language Model Unlearning
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Multi-Objective Large Language Model Unlearning
von: Pan, Zibin, et al.
Veröffentlicht: (2024)
von: Pan, Zibin, et al.
Veröffentlicht: (2024)
A Closer Look at Machine Unlearning for Large Language Models
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
Hierarchical Federated Unlearning for Large Language Models
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
Soft Prompting for Unlearning in Large Language Models
von: Bhaila, Karuna, et al.
Veröffentlicht: (2024)
von: Bhaila, Karuna, et al.
Veröffentlicht: (2024)
Large Language Model Unlearning via Embedding-Corrupted Prompts
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2024)
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
von: Li, Hongji, et al.
Veröffentlicht: (2025)
von: Li, Hongji, et al.
Veröffentlicht: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language Models
von: Entesari, Taha, et al.
Veröffentlicht: (2025)
von: Entesari, Taha, et al.
Veröffentlicht: (2025)
Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
von: Gu, Kang, et al.
Veröffentlicht: (2024)
von: Gu, Kang, et al.
Veröffentlicht: (2024)
Optimizing Diversity and Quality through Base-Aligned Model Collaboration
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
von: Wang, Yichen, et al.
Veröffentlicht: (2025)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024) -
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025) -
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
von: Huang, James Y., et al.
Veröffentlicht: (2025) -
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025) -
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)