Norm Anchors Make Model Edits Last

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Mingda, Zhu, Zhenghan, Miao, Ze'an, Fujisawa, Katsuki
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913092759191552
author Liu, Mingda
Zhu, Zhenghan
Miao, Ze'an
Fujisawa, Katsuki
author_facet Liu, Mingda
Zhu, Zhenghan
Miao, Ze'an
Fujisawa, Katsuki
contents Sequential Locate-and-Edit (L&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive norm-feedback loop, in which solved value vectors and edited MLP weights progressively amplify each other, degrading edit quality and eventually collapsing model capabilities. Our analysis shows that this feedback can yield approximately exponential norm growth under standard L&E dynamics, and can remain unresolved by existing increment-level regularizers or update clamps. We propose Norm-Anchor Scaling (NAS), a plug-in stabilizer that breaks this loop by rescaling each solved value vector to an original-model reference norm. Across multiple LLM backbones, datasets, and L&E editors, NAS extends the usable editing horizon by more than 4x and improves long-run editing performance by 72.2% on average, while preserving single-edit efficacy, with only a one-line modification and negligible computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Norm Anchors Make Model Edits Last
Liu, Mingda
Zhu, Zhenghan
Miao, Ze'an
Fujisawa, Katsuki
Machine Learning
Artificial Intelligence
Sequential Locate-and-Edit (L&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive norm-feedback loop, in which solved value vectors and edited MLP weights progressively amplify each other, degrading edit quality and eventually collapsing model capabilities. Our analysis shows that this feedback can yield approximately exponential norm growth under standard L&E dynamics, and can remain unresolved by existing increment-level regularizers or update clamps. We propose Norm-Anchor Scaling (NAS), a plug-in stabilizer that breaks this loop by rescaling each solved value vector to an original-model reference norm. Across multiple LLM backbones, datasets, and L&E editors, NAS extends the usable editing horizon by more than 4x and improves long-run editing performance by 72.2% on average, while preserving single-edit efficacy, with only a one-line modification and negligible computational overhead.
title Norm Anchors Make Model Edits Last
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.02543