Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thaker, Pratiksha, Hu, Shengyuan, Kale, Neil, Maurya, Yash, Wu, Zhiwei Steven, Smith, Virginia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Guardrail Baselines for Unlearning in LLMs
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
von: Hu, Shengyuan, et al.
Veröffentlicht: (2025)
von: Hu, Shengyuan, et al.
Veröffentlicht: (2025)
Membership Inference Attacks for Unseen Classes
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2025)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2025)
On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2023)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2023)
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
von: Hu, Shengyuan, et al.
Veröffentlicht: (2024)
von: Hu, Shengyuan, et al.
Veröffentlicht: (2024)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
von: Pang, Qi, et al.
Veröffentlicht: (2024)
von: Pang, Qi, et al.
Veröffentlicht: (2024)
Enhancing One-run Privacy Auditing with Quantile Regression-Based Membership Inference
von: Liu, Terrance, et al.
Veröffentlicht: (2025)
von: Liu, Terrance, et al.
Veröffentlicht: (2025)
PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
von: Kolawole, Steven, et al.
Veröffentlicht: (2025)
von: Kolawole, Steven, et al.
Veröffentlicht: (2025)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
Semantic Agreement Enables Efficient Open-Ended LLM Cascades
von: Soiffer, Duncan, et al.
Veröffentlicht: (2025)
von: Soiffer, Duncan, et al.
Veröffentlicht: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
Generate-then-Verify: Reconstructing Data from Limited Published Statistics
von: Liu, Terrance, et al.
Veröffentlicht: (2025)
von: Liu, Terrance, et al.
Veröffentlicht: (2025)
AI Governance and Accountability: An Analysis of Anthropic's Claude
von: Priyanshu, Aman, et al.
Veröffentlicht: (2024)
von: Priyanshu, Aman, et al.
Veröffentlicht: (2024)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
Measuring the Depth of LLM Unlearning via Activation Patching
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
von: Thaker, Khushboo, et al.
Veröffentlicht: (2025)
von: Thaker, Khushboo, et al.
Veröffentlicht: (2025)
LLM Unlearning with LLM Beliefs
von: Li, Kemou, et al.
Veröffentlicht: (2025)
von: Li, Kemou, et al.
Veröffentlicht: (2025)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
von: Sawant, Yash Ganpat
Veröffentlicht: (2026)
von: Sawant, Yash Ganpat
Veröffentlicht: (2026)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
von: Goel, Yash, et al.
Veröffentlicht: (2025)
von: Goel, Yash, et al.
Veröffentlicht: (2025)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
von: Kale, Sahil
Veröffentlicht: (2025)
von: Kale, Sahil
Veröffentlicht: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
LLM Unlearning Should Be Form-Independent
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
von: Ren, Richard, et al.
Veröffentlicht: (2024)
von: Ren, Richard, et al.
Veröffentlicht: (2024)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
von: Ye, Xiaotian, et al.
Veröffentlicht: (2026)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2026)
Representation-Guided Parameter-Efficient LLM Unlearning
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
von: Qian, Qi, et al.
Veröffentlicht: (2026)
von: Qian, Qi, et al.
Veröffentlicht: (2026)
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
von: Zhu, Wenhong, et al.
Veröffentlicht: (2024)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2024)
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
SEPS: A Separability Measure for Robust Unlearning in LLMs
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Guardrail Baselines for Unlearning in LLMs
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024) -
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
von: Hu, Shengyuan, et al.
Veröffentlicht: (2025) -
Membership Inference Attacks for Unseen Classes
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2025) -
On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2023) -
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
von: Hu, Shengyuan, et al.
Veröffentlicht: (2024)