SAKE: Steering Activations for Knowledge Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Scialanga, Marco, Laugel, Thibault, Grari, Vincent, Detyniecki, Marcin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913963288035328
author Scialanga, Marco
Laugel, Thibault
Grari, Vincent
Detyniecki, Marcin
author_facet Scialanga, Marco
Laugel, Thibault
Grari, Vincent
Detyniecki, Marcin
contents As Large Langue Models have been shown to memorize real-world facts, the need to update this knowledge in a controlled and efficient manner arises. Designed with these constraints in mind, Knowledge Editing (KE) approaches propose to alter specific facts in pretrained models. However, they have been shown to suffer from several limitations, including their lack of contextual robustness and their failure to generalize to logical implications related to the fact. To overcome these issues, we propose SAKE, a steering activation method that models a fact to be edited as a distribution rather than a single prompt. Leveraging Optimal Transport, SAKE alters the LLM behavior over a whole fact-related distribution, defined as paraphrases and logical implications. Several numerical experiments demonstrate the effectiveness of this method: SAKE is thus able to perform more robust edits than its existing counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAKE: Steering Activations for Knowledge Editing
Scialanga, Marco
Laugel, Thibault
Grari, Vincent
Detyniecki, Marcin
Artificial Intelligence
Computation and Language
Machine Learning
As Large Langue Models have been shown to memorize real-world facts, the need to update this knowledge in a controlled and efficient manner arises. Designed with these constraints in mind, Knowledge Editing (KE) approaches propose to alter specific facts in pretrained models. However, they have been shown to suffer from several limitations, including their lack of contextual robustness and their failure to generalize to logical implications related to the fact. To overcome these issues, we propose SAKE, a steering activation method that models a fact to be edited as a distribution rather than a single prompt. Leveraging Optimal Transport, SAKE alters the LLM behavior over a whole fact-related distribution, defined as paraphrases and logical implications. Several numerical experiments demonstrate the effectiveness of this method: SAKE is thus able to perform more robust edits than its existing counterparts.
title SAKE: Steering Activations for Knowledge Editing
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.01751