Fine-tuning Done Right in Model Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Wanli, Tang, Rui, Zang, Hongyu, Su, Du, Cao, Qi, Wang, Jingang, Shen, Huawei, Cheng, Xueqi, Sun, Fei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912928338280448
author Yang, Wanli
Tang, Rui
Zang, Hongyu
Su, Du
Cao, Qi
Wang, Jingang
Shen, Huawei
Cheng, Xueqi
Sun, Fei
author_facet Yang, Wanli
Tang, Rui
Zang, Hongyu
Su, Du
Cao, Qi
Wang, Jingang
Shen, Huawei
Cheng, Xueqi
Sun, Fei
contents Fine-tuning, a foundational method for adapting large language models, has long been considered ineffective for model editing. Here, we challenge this belief, arguing that the reported failure arises not from the inherent limitation of fine-tuning itself, but from adapting it to the sequential nature of the editing task, a single-pass depth-first pipeline that optimizes each sample to convergence before moving on. While intuitive, this depth-first pipeline coupled with sample-wise updating over-optimizes each edit and induces interference across edits. Our controlled experiments reveal that simply restoring fine-tuning to the standard breadth-first (i.e., epoch-based) pipeline with mini-batch optimization substantially improves its effectiveness for model editing. Moreover, fine-tuning in editing also suffers from suboptimal tuning parameter locations inherited from prior methods. Through systematic analysis of tuning locations, we derive LocFT-BF, a simple and effective localized editing method built on the restored fine-tuning framework. Extensive experiments across diverse LLMs and datasets demonstrate that LocFT-BF outperforms state-of-the-art methods by large margins. Notably, to our knowledge, it is the first to sustain 100K edits and 72B-parameter models,10 x beyond prior practice, without sacrificing general capabilities. By clarifying a long-standing misconception and introducing a principled localized tuning strategy, we advance fine-tuning from an underestimated baseline to a leading method for model editing, establishing a solid foundation for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-tuning Done Right in Model Editing
Yang, Wanli
Tang, Rui
Zang, Hongyu
Su, Du
Cao, Qi
Wang, Jingang
Shen, Huawei
Cheng, Xueqi
Sun, Fei
Computation and Language
Fine-tuning, a foundational method for adapting large language models, has long been considered ineffective for model editing. Here, we challenge this belief, arguing that the reported failure arises not from the inherent limitation of fine-tuning itself, but from adapting it to the sequential nature of the editing task, a single-pass depth-first pipeline that optimizes each sample to convergence before moving on. While intuitive, this depth-first pipeline coupled with sample-wise updating over-optimizes each edit and induces interference across edits. Our controlled experiments reveal that simply restoring fine-tuning to the standard breadth-first (i.e., epoch-based) pipeline with mini-batch optimization substantially improves its effectiveness for model editing. Moreover, fine-tuning in editing also suffers from suboptimal tuning parameter locations inherited from prior methods. Through systematic analysis of tuning locations, we derive LocFT-BF, a simple and effective localized editing method built on the restored fine-tuning framework. Extensive experiments across diverse LLMs and datasets demonstrate that LocFT-BF outperforms state-of-the-art methods by large margins. Notably, to our knowledge, it is the first to sustain 100K edits and 72B-parameter models,10 x beyond prior practice, without sacrificing general capabilities. By clarifying a long-standing misconception and introducing a principled localized tuning strategy, we advance fine-tuning from an underestimated baseline to a leading method for model editing, establishing a solid foundation for future research.
title Fine-tuning Done Right in Model Editing
topic Computation and Language
url https://arxiv.org/abs/2509.22072