Representation-Guided Parameter-Efficient LLM Unlearning
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiao, Zeguan, Mo, Lang, Chen, Yun, Yang, Lei, Zhao, Jiehui, Yang, Lili, Chen, Guanhua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
por: Xiao, Zeguan, et al.
Publicado: (2026)
por: Xiao, Zeguan, et al.
Publicado: (2026)
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
por: Xiao, Zeguan, et al.
Publicado: (2026)
por: Xiao, Zeguan, et al.
Publicado: (2026)
Distract Large Language Models for Automatic Jailbreak Attack
por: Xiao, Zeguan, et al.
Publicado: (2024)
por: Xiao, Zeguan, et al.
Publicado: (2024)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
por: Xiao, Zeguan, et al.
Publicado: (2025)
por: Xiao, Zeguan, et al.
Publicado: (2025)
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
por: Xiao, Zeguan, et al.
Publicado: (2025)
por: Xiao, Zeguan, et al.
Publicado: (2025)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
por: Wang, Hanqing, et al.
Publicado: (2024)
por: Wang, Hanqing, et al.
Publicado: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
por: Yang, Yan, et al.
Publicado: (2024)
por: Yang, Yan, et al.
Publicado: (2024)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
por: Wu, Wanxing, et al.
Publicado: (2026)
por: Wu, Wanxing, et al.
Publicado: (2026)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
por: Lai, Peng, et al.
Publicado: (2025)
por: Lai, Peng, et al.
Publicado: (2025)
Toward Automated Robustness Evaluation of Mathematical Reasoning
por: Hou, Yutao, et al.
Publicado: (2025)
por: Hou, Yutao, et al.
Publicado: (2025)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
por: Su, Junyou, et al.
Publicado: (2026)
por: Su, Junyou, et al.
Publicado: (2026)
G2: Guided Generation for Enhanced Output Diversity in LLMs
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
por: Liu, Yezi, et al.
Publicado: (2025)
por: Liu, Yezi, et al.
Publicado: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
por: Hou, Yutao, et al.
Publicado: (2026)
por: Hou, Yutao, et al.
Publicado: (2026)
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
por: Zhao, Guoshenghui, et al.
Publicado: (2025)
por: Zhao, Guoshenghui, et al.
Publicado: (2025)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
por: Liu, Yezi, et al.
Publicado: (2025)
por: Liu, Yezi, et al.
Publicado: (2025)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
por: Liu, Dongxu, et al.
Publicado: (2024)
por: Liu, Dongxu, et al.
Publicado: (2024)
Pi-SQL: Enhancing Text-to-SQL with Fine-Grained Guidance from Pivot Programming Languages
por: chi, Yongdong, et al.
Publicado: (2025)
por: chi, Yongdong, et al.
Publicado: (2025)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
por: Ruan, Zhiwen, et al.
Publicado: (2026)
por: Ruan, Zhiwen, et al.
Publicado: (2026)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
por: Chen, Hao, et al.
Publicado: (2026)
por: Chen, Hao, et al.
Publicado: (2026)
CLUE: Conflict-guided Localization for LLM Unlearning Framework
por: Chen, Hang, et al.
Publicado: (2025)
por: Chen, Hang, et al.
Publicado: (2025)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
por: Sondej, Filip, et al.
Publicado: (2025)
por: Sondej, Filip, et al.
Publicado: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
por: Lai, Peng, et al.
Publicado: (2026)
por: Lai, Peng, et al.
Publicado: (2026)
Towards Robust and Parameter-Efficient Knowledge Unlearning for LLMs
por: Cha, Sungmin, et al.
Publicado: (2024)
por: Cha, Sungmin, et al.
Publicado: (2024)
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
por: Zhang, Hengxiang, et al.
Publicado: (2024)
por: Zhang, Hengxiang, et al.
Publicado: (2024)
Unlocking Multi-View Insights in Knowledge-Dense Retrieval-Augmented Generation
por: Chen, Guanhua, et al.
Publicado: (2024)
por: Chen, Guanhua, et al.
Publicado: (2024)
Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
por: Yang, Yan, et al.
Publicado: (2025)
por: Yang, Yan, et al.
Publicado: (2025)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
por: Xue, Tianci, et al.
Publicado: (2023)
por: Xue, Tianci, et al.
Publicado: (2023)
Parameter-Efficient Fine-Tuning With Adapters
por: Chen, Keyu, et al.
Publicado: (2024)
por: Chen, Keyu, et al.
Publicado: (2024)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
por: Dorna, Vineeth, et al.
Publicado: (2025)
por: Dorna, Vineeth, et al.
Publicado: (2025)
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
por: Li, Yixia, et al.
Publicado: (2024)
por: Li, Yixia, et al.
Publicado: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
por: Sondej, Filip, et al.
Publicado: (2025)
por: Sondej, Filip, et al.
Publicado: (2025)
MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models
por: Luo, Tongxu, et al.
Publicado: (2024)
por: Luo, Tongxu, et al.
Publicado: (2024)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
por: Yamashita, Tomoya, et al.
Publicado: (2025)
por: Yamashita, Tomoya, et al.
Publicado: (2025)
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
por: Zhai, Naixin, et al.
Publicado: (2026)
por: Zhai, Naixin, et al.
Publicado: (2026)
CATNIP: LLM Unlearning via Calibrated and Tokenized Negative Preference Alignment
por: Yang, Zhengbang, et al.
Publicado: (2026)
por: Yang, Zhengbang, et al.
Publicado: (2026)
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
por: Zhang, Xiang, et al.
Publicado: (2025)
por: Zhang, Xiang, et al.
Publicado: (2025)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
por: Ji, Jiabao, et al.
Publicado: (2024)
por: Ji, Jiabao, et al.
Publicado: (2024)
Ejemplares similares
-
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
por: Xiao, Zeguan, et al.
Publicado: (2026) -
Modeling LLM Unlearning as an Asymmetric Two-Task Learning Problem
por: Xiao, Zeguan, et al.
Publicado: (2026) -
Distract Large Language Models for Automatic Jailbreak Attack
por: Xiao, Zeguan, et al.
Publicado: (2024) -
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
por: Xiao, Zeguan, et al.
Publicado: (2025) -
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
por: Xiao, Zeguan, et al.
Publicado: (2025)