Measuring the Depth of LLM Unlearning via Activation Patching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jaeung, Kim, Dohyun, Jo, Jaemin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
von: Poonia, Ansh, et al.
Veröffentlicht: (2025)
von: Poonia, Ansh, et al.
Veröffentlicht: (2025)
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
von: Kim, Jaemin, et al.
Veröffentlicht: (2026)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG
von: Lee, Changmin, et al.
Veröffentlicht: (2026)
von: Lee, Changmin, et al.
Veröffentlicht: (2026)
LLM Unlearning Without an Expert Curated Dataset
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
Unlearning Comparator: A Visual Analytics System for Comparative Evaluation of Machine Unlearning Methods
von: Lee, Jaeung, et al.
Veröffentlicht: (2025)
von: Lee, Jaeung, et al.
Veröffentlicht: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
Suppression or Deletion: A Restoration-Based Representation-Level Analysis of Machine Unlearning
von: Jang, Yurim, et al.
Veröffentlicht: (2026)
von: Jang, Yurim, et al.
Veröffentlicht: (2026)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
von: Zala, Abhay, et al.
Veröffentlicht: (2023)
von: Zala, Abhay, et al.
Veröffentlicht: (2023)
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
von: Kim, Dahyun, et al.
Veröffentlicht: (2023)
von: Kim, Dahyun, et al.
Veröffentlicht: (2023)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
von: Lin, Han, et al.
Veröffentlicht: (2023)
von: Lin, Han, et al.
Veröffentlicht: (2023)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
von: Zala, Abhay, et al.
Veröffentlicht: (2024)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
von: Kim, Jeonghye, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2025)
DepthCharge: A Domain-Agnostic Framework for Measuring Depth-Dependent Knowledge in Large Language Models
von: Sheppert, Alexander
Veröffentlicht: (2026)
von: Sheppert, Alexander
Veröffentlicht: (2026)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
Turning LLM Activations Quantization-Friendly
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
Large Language Model Unlearning via Embedding-Corrupted Prompts
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Geometric-disentangelment Unlearning
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Assessing LLM Reasoning Steps via Principal Knowledge Grounding
von: Hwang, Hyeon, et al.
Veröffentlicht: (2025)
von: Hwang, Hyeon, et al.
Veröffentlicht: (2025)
PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2024)
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching
von: Poonia, Ansh, et al.
Veröffentlicht: (2025) -
Adaptive Guidance for Retrieval-Augmented Masked Diffusion Models
von: Kim, Jaemin, et al.
Veröffentlicht: (2026) -
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024) -
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
von: Sondej, Filip, et al.
Veröffentlicht: (2025) -
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction
von: Lee, Yooseop, et al.
Veröffentlicht: (2025)