Representation-Guided Parameter-Efficient LLM Unlearning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Zeguan, Mo, Lang, Chen, Yun, Yang, Lei, Zhao, Jiehui, Yang, Lili, Chen, Guanhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026)
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
von: Wang, Hanqing, et al.
Veröffentlicht: (2024)
von: Wang, Hanqing, et al.
Veröffentlicht: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
von: Wu, Wanxing, et al.
Veröffentlicht: (2026)
von: Wu, Wanxing, et al.
Veröffentlicht: (2026)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
von: Lai, Peng, et al.
Veröffentlicht: (2025)
von: Lai, Peng, et al.
Veröffentlicht: (2025)
Toward Automated Robustness Evaluation of Mathematical Reasoning
von: Hou, Yutao, et al.
Veröffentlicht: (2025)
von: Hou, Yutao, et al.
Veröffentlicht: (2025)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026)
von: Su, Junyou, et al.
Veröffentlicht: (2026)
G2: Guided Generation for Enhanced Output Diversity in LLMs
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
von: Liu, Dongxu, et al.
Veröffentlicht: (2024)
von: Liu, Dongxu, et al.
Veröffentlicht: (2024)
Pi-SQL: Enhancing Text-to-SQL with Fine-Grained Guidance from Pivot Programming Languages
von: chi, Yongdong, et al.
Veröffentlicht: (2025)
von: chi, Yongdong, et al.
Veröffentlicht: (2025)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
CLUE: Conflict-guided Localization for LLM Unlearning Framework
von: Chen, Hang, et al.
Veröffentlicht: (2025)
von: Chen, Hang, et al.
Veröffentlicht: (2025)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
von: Lai, Peng, et al.
Veröffentlicht: (2026)
von: Lai, Peng, et al.
Veröffentlicht: (2026)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
von: Cha, Sungmin, et al.
Veröffentlicht: (2024)
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
Unlocking Multi-View Insights in Knowledge-Dense Retrieval-Augmented Generation
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
Parameter-Efficient Fine-Tuning With Adapters
von: Chen, Keyu, et al.
Veröffentlicht: (2024)
von: Chen, Keyu, et al.
Veröffentlicht: (2024)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
von: Dorna, Vineeth, et al.
Veröffentlicht: (2025)
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
von: Li, Yixia, et al.
Veröffentlicht: (2024)
von: Li, Yixia, et al.
Veröffentlicht: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models
von: Luo, Tongxu, et al.
Veröffentlicht: (2024)
von: Luo, Tongxu, et al.
Veröffentlicht: (2024)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
von: Zhai, Naixin, et al.
Veröffentlicht: (2026)
von: Zhai, Naixin, et al.
Veröffentlicht: (2026)
CATNIP: LLM Unlearning via Calibrated and Tokenized Negative Preference Alignment
von: Yang, Zhengbang, et al.
Veröffentlicht: (2026)
von: Yang, Zhengbang, et al.
Veröffentlicht: (2026)
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
von: Zhang, Xiang, et al.
Veröffentlicht: (2025)
von: Zhang, Xiang, et al.
Veröffentlicht: (2025)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026) -
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
von: Xiao, Zeguan, et al.
Veröffentlicht: (2026) -
Distract Large Language Models for Automatic Jailbreak Attack
von: Xiao, Zeguan, et al.
Veröffentlicht: (2024) -
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025) -
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)