Existing Large Language Model Unlearning Evaluations Are Inconclusive
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Zhili, Xu, Yixuan Even, Robey, Alexander, Kirk, Robert, Davies, Xander, Gal, Yarin, Schwarzschild, Avi, Kolter, J. Zico |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Antidistillation Sampling
by: Savani, Yash, et al.
Published: (2025)
by: Savani, Yash, et al.
Published: (2025)
TOFU: A Task of Fictitious Unlearning for LLMs
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024)
by: Schwarzschild, Avi, et al.
Published: (2024)
Compressed Sensing for Capability Localization in Large Language Models
by: Bair, Anna, et al.
Published: (2026)
by: Bair, Anna, et al.
Published: (2026)
Forcing Diffuse Distributions out of Language Models
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Antidistillation Fingerprinting
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Base Models Look Human To AI Detectors
by: Xu, Yixuan Even, et al.
Published: (2026)
by: Xu, Yixuan Even, et al.
Published: (2026)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Idiosyncrasies in Large Language Models
by: Sun, Mingjie, et al.
Published: (2025)
by: Sun, Mingjie, et al.
Published: (2025)
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)
by: Zou, Andy, et al.
Published: (2025)
Massive Activations in Large Language Models
by: Sun, Mingjie, et al.
Published: (2024)
by: Sun, Mingjie, et al.
Published: (2024)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
by: Dorna, Vineeth, et al.
Published: (2025)
by: Dorna, Vineeth, et al.
Published: (2025)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Language Models Change Facts Based on the Way You Talk
by: Kearney, Matthew, et al.
Published: (2025)
by: Kearney, Matthew, et al.
Published: (2025)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
An Axiomatic Approach to Model-Agnostic Concept Explanations
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
When Should We Introduce Safety Interventions During Pretraining?
by: Sam, Dylan, et al.
Published: (2026)
by: Sam, Dylan, et al.
Published: (2026)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
Evaluating Deep Unlearning in Large Language Models
by: Wu, Ruihan, et al.
Published: (2024)
by: Wu, Ruihan, et al.
Published: (2024)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Unnatural Languages Are Not Bugs but Features for LLMs
by: Duan, Keyu, et al.
Published: (2025)
by: Duan, Keyu, et al.
Published: (2025)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Large Language Model Unlearning
by: Yao, Yuanshun, et al.
Published: (2023)
by: Yao, Yuanshun, et al.
Published: (2023)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
Extrapolation by Association: Length Generalization Transfer in Transformers
by: Cai, Ziyang, et al.
Published: (2025)
by: Cai, Ziyang, et al.
Published: (2025)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
by: Li, Maximilian, et al.
Published: (2023)
by: Li, Maximilian, et al.
Published: (2023)
Bayesian Preference Elicitation with Language Models
by: Handa, Kunal, et al.
Published: (2024)
by: Handa, Kunal, et al.
Published: (2024)
Similar Items
-
Antidistillation Sampling
by: Savani, Yash, et al.
Published: (2025) -
TOFU: A Task of Fictitious Unlearning for LLMs
by: Maini, Pratyush, et al.
Published: (2024) -
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
by: Ball, Sarah, et al.
Published: (2025) -
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024) -
Compressed Sensing for Capability Localization in Large Language Models
by: Bair, Anna, et al.
Published: (2026)