Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
Fuente:
arXiv
Saved in:
| Main Authors: | Lang, Yicheng, Guo, Kehan, Huang, Yue, Zhou, Yujun, Zhuang, Haomin, Yang, Tianyu, Su, Yao, Zhang, Xiangliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026)
by: Wang, Xiangqi, et al.
Published: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024)
by: Zhou, Yujun, et al.
Published: (2024)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
Reliable Control-Point Selection for Steering Reasoning in Large Language Models
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Exploring Multi-Temperature Strategies for Token- and Rollout-Level Control in RLVR
by: Zhuang, Haomin, et al.
Published: (2025)
by: Zhuang, Haomin, et al.
Published: (2025)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
AIRGuard: Guarding Agent Actions with Runtime Authority Control
by: Qin, Suliu, et al.
Published: (2026)
by: Qin, Suliu, et al.
Published: (2026)
Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
by: Huang, Yue, et al.
Published: (2024)
by: Huang, Yue, et al.
Published: (2024)
AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
AgentClick: A Skill-Based Human-in-the-Loop Review Layer for Terminal AI Agents
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond
by: Guo, Kehan, et al.
Published: (2025)
by: Guo, Kehan, et al.
Published: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
by: Xu, Zixiang, et al.
Published: (2025)
by: Xu, Zixiang, et al.
Published: (2025)
ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
SkillGen: Verified Inference-Time Agent Skill Synthesis
by: Ma, Yuchen, et al.
Published: (2026)
by: Ma, Yuchen, et al.
Published: (2026)
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
by: Shah, Raj Sanjay, et al.
Published: (2026)
by: Shah, Raj Sanjay, et al.
Published: (2026)
Cognitive Echo: Enhancing think‐aloud protocols with LLM ‐based simulated students
by: Longwei Zheng, et al.
Published: (2025)
by: Longwei Zheng, et al.
Published: (2025)
Discrete Mean Field Games on Finite Graphs as Initial Value Optimization
by: Feng, Yaxin, et al.
Published: (2026)
by: Feng, Yaxin, et al.
Published: (2026)
Capability-Oriented Training Induced Alignment Risk
by: Zhou, Yujun, et al.
Published: (2026)
by: Zhou, Yujun, et al.
Published: (2026)
PrivacyCD: Hierarchical Unlearning for Protecting Student Privacy in Cognitive Diagnosis
by: Hou, Mingliang, et al.
Published: (2025)
by: Hou, Mingliang, et al.
Published: (2025)
LMCD: Language Models are Zeroshot Cognitive Diagnosis Learners
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
Quasiparticle Interference Kernel Extraction with Variational Autoencoders via Latent Alignment
by: Ji, Yingshuai, et al.
Published: (2025)
by: Ji, Yingshuai, et al.
Published: (2025)
LLM Unlearning with LLM Beliefs
by: Li, Kemou, et al.
Published: (2025)
by: Li, Kemou, et al.
Published: (2025)
ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
by: Guo, Taicheng, et al.
Published: (2023)
by: Guo, Taicheng, et al.
Published: (2023)
Blunted P300 Prospectively Bridges Cognitive Reappraisal and Depressive Symptoms
by: Kehan Li, et al.
Published: (2025)
by: Kehan Li, et al.
Published: (2025)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
by: Dorna, Vineeth, et al.
Published: (2025)
by: Dorna, Vineeth, et al.
Published: (2025)
Representation-Guided Parameter-Efficient LLM Unlearning
by: Xiao, Zeguan, et al.
Published: (2026)
by: Xiao, Zeguan, et al.
Published: (2026)
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
by: Yang, Puning, et al.
Published: (2025)
by: Yang, Puning, et al.
Published: (2025)
AI Alignment Breaks at the Edge
by: Bao, Han, et al.
Published: (2026)
by: Bao, Han, et al.
Published: (2026)
Similar Items
-
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026) -
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024) -
Defending Jailbreak Prompts via In-Context Adversarial Game
by: Zhou, Yujun, et al.
Published: (2024) -
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025) -
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)