Eight Methods to Evaluate Robust Unlearning in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lynch, Aengus, Guo, Phillip, Ewart, Aidan, Casper, Stephen, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Diverse Preference Learning for Capabilities and Alignment
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Unlearning in LLMs: Methods, Evaluation, and Open Challenges
von: Lizzo, Tyler, et al.
Veröffentlicht: (2026)
von: Lizzo, Tyler, et al.
Veröffentlicht: (2026)
Layered Unlearning for Adversarial Relearning
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
SEPS: A Separability Measure for Robust Unlearning in LLMs
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
von: Jeung, Wonje, et al.
Veröffentlicht: (2025)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
The Persistent Vulnerability of Aligned AI Systems
von: Lynch, Aengus
Veröffentlicht: (2026)
von: Lynch, Aengus
Veröffentlicht: (2026)
Guardrail Baselines for Unlearning in LLMs
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
von: Thaker, Pratiksha, et al.
Veröffentlicht: (2024)
Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs
von: Bouchard, Dylan
Veröffentlicht: (2024)
von: Bouchard, Dylan
Veröffentlicht: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
Leverage Unlearning to Sanitize LLMs
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
von: Ford, Casey, et al.
Veröffentlicht: (2026)
von: Ford, Casey, et al.
Veröffentlicht: (2026)
Rethinking Machine Unlearning for Large Language Models
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
von: Rezkellah, Fatmazohra, et al.
Veröffentlicht: (2025)
von: Rezkellah, Fatmazohra, et al.
Veröffentlicht: (2025)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
von: Kirch, Nathalie, et al.
Veröffentlicht: (2024)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2024)
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
von: Hwang, Kyomin, et al.
Veröffentlicht: (2025)
von: Hwang, Kyomin, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
Improving LLM Unlearning Robustness via Random Perturbations
von: Huu-Tien, Dang, et al.
Veröffentlicht: (2025)
von: Huu-Tien, Dang, et al.
Veröffentlicht: (2025)
ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs
von: Chen, Xunlei, et al.
Veröffentlicht: (2026)
von: Chen, Xunlei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024) -
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025) -
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024) -
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025) -
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)