CAP: Controllable Alignment Prompting for Unlearning in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zhaokun, Guo, Jinyu, Pu, Jingwen, Pu, Hongli, Yang, Meng, Chen, Xunlei, Ou, Jie, Li, Wenyi, Luo, Guangchun, Tian, Wenhong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909045418360832
author Wang, Zhaokun
Guo, Jinyu
Pu, Jingwen
Pu, Hongli
Yang, Meng
Chen, Xunlei
Ou, Jie
Li, Wenyi
Luo, Guangchun
Tian, Wenhong
author_facet Wang, Zhaokun
Guo, Jinyu
Pu, Jingwen
Pu, Hongli
Yang, Meng
Chen, Xunlei
Ou, Jie
Li, Wenyi
Luo, Guangchun
Tian, Wenhong
contents Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-modifying methods face fundamental limitations: high computational costs, uncontrollable forgetting boundaries, and strict dependency on model weight access. These constraints render them impractical for closed-source models, yet current non-invasive alternatives remain unsystematic and reliant on empirical experience. To address these challenges, we propose the Controllable Alignment Prompting for Unlearning (CAP) framework, an end-to-end prompt-driven unlearning paradigm. CAP decouples unlearning into a learnable prompt optimization process via reinforcement learning, where a prompt generator collaborates with the LLM to suppress target knowledge while preserving general capabilities selectively. This approach enables reversible knowledge restoration through prompt revocation. Extensive experiments demonstrate that CAP achieves precise, controllable unlearning without updating model parameters, establishing a dynamic alignment mechanism that overcomes the transferability limitations of prior methods.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21251
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CAP: Controllable Alignment Prompting for Unlearning in LLMs
Wang, Zhaokun
Guo, Jinyu
Pu, Jingwen
Pu, Hongli
Yang, Meng
Chen, Xunlei
Ou, Jie
Li, Wenyi
Luo, Guangchun
Tian, Wenhong
Machine Learning
Artificial Intelligence
Large language models (LLMs) trained on unfiltered corpora inherently risk retaining sensitive information, necessitating selective knowledge unlearning for regulatory compliance and ethical safety. However, existing parameter-modifying methods face fundamental limitations: high computational costs, uncontrollable forgetting boundaries, and strict dependency on model weight access. These constraints render them impractical for closed-source models, yet current non-invasive alternatives remain unsystematic and reliant on empirical experience. To address these challenges, we propose the Controllable Alignment Prompting for Unlearning (CAP) framework, an end-to-end prompt-driven unlearning paradigm. CAP decouples unlearning into a learnable prompt optimization process via reinforcement learning, where a prompt generator collaborates with the LLM to suppress target knowledge while preserving general capabilities selectively. This approach enables reversible knowledge restoration through prompt revocation. Extensive experiments demonstrate that CAP achieves precise, controllable unlearning without updating model parameters, establishing a dynamic alignment mechanism that overcomes the transferability limitations of prior methods.
title CAP: Controllable Alignment Prompting for Unlearning in LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.21251