Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hang, Zhu, Jiaying, Chen, Hongyang, Liu, Hongxu, Yang, Xinyu, Wang, Wenya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914539466915840
author Chen, Hang
Zhu, Jiaying
Chen, Hongyang
Liu, Hongxu
Yang, Xinyu
Wang, Wenya
author_facet Chen, Hang
Zhu, Jiaying
Chen, Hongyang
Liu, Hongxu
Yang, Xinyu
Wang, Wenya
contents The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpretability for targeted parameter updates. However, this paradigm rests on a fundamental yet unverified assumption: can mechanisms derived from current static parameters reliably guide future dynamic parameter updates? To investigate this, we systematically track the structural evolution of Transformer circuits throughout the supervised fine-tuning (SFT) process, revealing the underlying dynamics of task mechanisms. We introduce three novel metrics-Circuit Distance, Circuit Stability, and Circuit Conflict-to analyze circuit evolution across three dimensions: neural migration, semantic stability, and cross-task interference. Our empirical results reveal that circuits inherently exhibit "Free Evolution" during parameter updates. Consequently, static mechanisms extracted from current states inevitably suffer from temporal latency, making them fundamentally inadequate for guiding future states. Moreover, by deconstructing the "illusion of effectiveness" in existing methods, this work underscores the necessity of "foresight" in mechanistic localization and proposes a predictive framework for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06076
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
Chen, Hang
Zhu, Jiaying
Chen, Hongyang
Liu, Hongxu
Yang, Xinyu
Wang, Wenya
Computation and Language
The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpretability for targeted parameter updates. However, this paradigm rests on a fundamental yet unverified assumption: can mechanisms derived from current static parameters reliably guide future dynamic parameter updates? To investigate this, we systematically track the structural evolution of Transformer circuits throughout the supervised fine-tuning (SFT) process, revealing the underlying dynamics of task mechanisms. We introduce three novel metrics-Circuit Distance, Circuit Stability, and Circuit Conflict-to analyze circuit evolution across three dimensions: neural migration, semantic stability, and cross-task interference. Our empirical results reveal that circuits inherently exhibit "Free Evolution" during parameter updates. Consequently, static mechanisms extracted from current states inevitably suffer from temporal latency, making them fundamentally inadequate for guiding future states. Moreover, by deconstructing the "illusion of effectiveness" in existing methods, this work underscores the necessity of "foresight" in mechanistic localization and proposes a predictive framework for future research.
title Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
topic Computation and Language
url https://arxiv.org/abs/2605.06076