SEPS: A Separability Measure for Robust Unlearning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jeung, Wonje, Yoon, Sangyeon, No, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
R-TOFU: Unlearning in Large Reasoning Models
by: Yoon, Sangyeon, et al.
Published: (2025)
by: Yoon, Sangyeon, et al.
Published: (2025)
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
DUSK: Do Not Unlearn Shared Knowledge
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
Adversarial Sample-Based Approach for Tighter Privacy Auditing in Final Model-Only Scenarios
by: Yoon, Sangyeon, et al.
Published: (2024)
by: Yoon, Sangyeon, et al.
Published: (2024)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
by: Hong, Hyesoo, et al.
Published: (2026)
by: Hong, Hyesoo, et al.
Published: (2026)
Large Language Models Still Exhibit Bias in Long Text
by: Jeung, Wonje, et al.
Published: (2024)
by: Jeung, Wonje, et al.
Published: (2024)
Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
by: Hwang, Kyomin, et al.
Published: (2026)
by: Hwang, Kyomin, et al.
Published: (2026)
An Information Theoretic Evaluation Metric For Strong Unlearning
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
Eight Methods to Evaluate Robust Unlearning in LLMs
by: Lynch, Aengus, et al.
Published: (2024)
by: Lynch, Aengus, et al.
Published: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2025)
by: Kim, Bumjun, et al.
Published: (2025)
Guardrail Baselines for Unlearning in LLMs
by: Thaker, Pratiksha, et al.
Published: (2024)
by: Thaker, Pratiksha, et al.
Published: (2024)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Leverage Unlearning to Sanitize LLMs
by: Boutet, Antoine, et al.
Published: (2025)
by: Boutet, Antoine, et al.
Published: (2025)
Unlearning in LLMs: Methods, Evaluation, and Open Challenges
by: Lizzo, Tyler, et al.
Published: (2026)
by: Lizzo, Tyler, et al.
Published: (2026)
TOFU: A Task of Fictitious Unlearning for LLMs
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Representation Bending for Large Language Model Safety
by: Yousefpour, Ashkan, et al.
Published: (2025)
by: Yousefpour, Ashkan, et al.
Published: (2025)
LLMs Can Infer Political Alignment from Online Conversations
by: Lee, Byunghwee, et al.
Published: (2026)
by: Lee, Byunghwee, et al.
Published: (2026)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
by: Hwang, Kyomin, et al.
Published: (2025)
by: Hwang, Kyomin, et al.
Published: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
by: Tutek, Martin, et al.
Published: (2025)
by: Tutek, Martin, et al.
Published: (2025)
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
by: Thaker, Pratiksha, et al.
Published: (2024)
by: Thaker, Pratiksha, et al.
Published: (2024)
Improving LLM Unlearning Robustness via Random Perturbations
by: Huu-Tien, Dang, et al.
Published: (2025)
by: Huu-Tien, Dang, et al.
Published: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
Less is More: Geometric Unlearning for LLMs with Minimal Data Disclosure
by: Tan, Chenchen, et al.
Published: (2026)
by: Tan, Chenchen, et al.
Published: (2026)
This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
Learning-Time Encoding Shapes Unlearning in LLMs
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
by: Hu, Xinshuo, et al.
Published: (2023)
by: Hu, Xinshuo, et al.
Published: (2023)
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
by: Luo, Yinyi, et al.
Published: (2025)
by: Luo, Yinyi, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Tool Unlearning for Tool-Augmented LLMs
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
Similar Items
-
R-TOFU: Unlearning in Large Reasoning Models
by: Yoon, Sangyeon, et al.
Published: (2025) -
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
by: Jeung, Wonje, et al.
Published: (2025) -
DUSK: Do Not Unlearn Shared Knowledge
by: Jeung, Wonje, et al.
Published: (2025) -
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
by: Yoon, Sangyeon, et al.
Published: (2026) -
Adversarial Sample-Based Approach for Tighter Privacy Auditing in Final Model-Only Scenarios
by: Yoon, Sangyeon, et al.
Published: (2024)