Probing Knowledge Holes in Unlearned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ko, Myeongseob, Just, Hoang Anh, Fleming, Charles, Jin, Ming, Jia, Ruoxi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
DiPT: Enhancing LLM reasoning through diversified perspective-taking
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Characterizing Model-Native Skills
by: Kang, Feiyang, et al.
Published: (2026)
by: Kang, Feiyang, et al.
Published: (2026)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward Passes
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
by: Sel, Bilgehan, et al.
Published: (2024)
by: Sel, Bilgehan, et al.
Published: (2024)
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
by: Ko, Myeongseob, et al.
Published: (2026)
by: Ko, Myeongseob, et al.
Published: (2026)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Injecting Measurement Information Yields a Fast and Noise-Robust Diffusion-Based Inverse Problem Solver
by: Patsenker, Jonathan, et al.
Published: (2025)
by: Patsenker, Jonathan, et al.
Published: (2025)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
SoK: Unlearnability and Unlearning for Model Dememorization
by: Zhang, Mengying, et al.
Published: (2026)
by: Zhang, Mengying, et al.
Published: (2026)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
by: Wang, Zhaokun, et al.
Published: (2026)
by: Wang, Zhaokun, et al.
Published: (2026)
ROKA: Robust Knowledge Unlearning against Adversaries
by: Shin, Jinmyeong, et al.
Published: (2026)
by: Shin, Jinmyeong, et al.
Published: (2026)
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
by: Mishra, Himanshu, et al.
Published: (2026)
by: Mishra, Himanshu, et al.
Published: (2026)
Efficient Knowledge Graph Unlearning with Zeroth-order Information
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Federated Knowledge Graph Unlearning via Diffusion Model
by: Liu, Bingchen, et al.
Published: (2024)
by: Liu, Bingchen, et al.
Published: (2024)
Machine Unlearning: Solutions and Challenges
by: Xu, Jie, et al.
Published: (2023)
by: Xu, Jie, et al.
Published: (2023)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
Surgical Knowledge Rewrite in Compact LLMs: An 'Unlearn-then-Learn' Strategy with ($IA^3$) for Localized Factual Modulation and Catastrophic Forgetting Mitigation
by: Ngugi, Stanley
Published: (2025)
by: Ngugi, Stanley
Published: (2025)
Debiasing Machine Unlearning with Counterfactual Examples
by: Chen, Ziheng, et al.
Published: (2024)
by: Chen, Ziheng, et al.
Published: (2024)
Hessian-Free Online Certified Unlearning
by: Qiao, Xinbao, et al.
Published: (2024)
by: Qiao, Xinbao, et al.
Published: (2024)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
by: Polowczyk, Agnieszka, et al.
Published: (2025)
by: Polowczyk, Agnieszka, et al.
Published: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
by: Lin, Yujie, et al.
Published: (2026)
by: Lin, Yujie, et al.
Published: (2026)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Understanding and Preserving Safety in Fine-Tuned LLMs
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025)
by: Scholten, Yan, et al.
Published: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
by: Mushtaq, Erum, et al.
Published: (2025)
by: Mushtaq, Erum, et al.
Published: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Data Valuation and Selection in a Federated Model Marketplace
by: Li, Wenqian, et al.
Published: (2025)
by: Li, Wenqian, et al.
Published: (2025)
The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed Samples
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
Efficient Bilevel Optimization for Meta Label Correction in Noisy Label Learning
by: Nguyen, Ba Hoang Anh, et al.
Published: (2026)
by: Nguyen, Ba Hoang Anh, et al.
Published: (2026)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Similar Items
-
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025) -
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
by: Ko, Myeongseob, et al.
Published: (2024) -
DiPT: Enhancing LLM reasoning through diversified perspective-taking
by: Just, Hoang Anh, et al.
Published: (2024) -
Characterizing Model-Native Skills
by: Kang, Feiyang, et al.
Published: (2026) -
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)