Saved in:
Bibliographic Details
Main Authors: Kangaslahti, Sara, Rosenfeld, Elan, Saphra, Naomi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.15872
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915877004247040
author Kangaslahti, Sara
Rosenfeld, Elan
Saphra, Naomi
author_facet Kangaslahti, Sara
Rosenfeld, Elan
Saphra, Naomi
contents Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. Studying these breakthroughs enables a deeper understanding of learning dynamics, but only when they are properly identified. This paper argues that similar breakthroughs occur frequently throughout training but they are obscured by a loss metric that collapses all variation into a single scalar. To find these hidden transitions, we introduce POLCA, a method for decomposing changes in loss along arbitrary bases of the low-rank training subspace. We use our method to identify clusters of samples that share similar changes in loss during training, disaggregating the overall loss into that of smaller groups of conceptually similar data. We validate our method on synthetic arithmetic and natural language tasks, showing that POLCA recovers clusters that represent interpretable breakthroughs in the model's capabilities. We demonstrate the promise of these hidden phase transitions as a tool for unsupervised interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15872
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hidden Breakthroughs in Language Model Training
Kangaslahti, Sara
Rosenfeld, Elan
Saphra, Naomi
Machine Learning
Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. Studying these breakthroughs enables a deeper understanding of learning dynamics, but only when they are properly identified. This paper argues that similar breakthroughs occur frequently throughout training but they are obscured by a loss metric that collapses all variation into a single scalar. To find these hidden transitions, we introduce POLCA, a method for decomposing changes in loss along arbitrary bases of the low-rank training subspace. We use our method to identify clusters of samples that share similar changes in loss during training, disaggregating the overall loss into that of smaller groups of conceptually similar data. We validate our method on synthetic arithmetic and natural language tasks, showing that POLCA recovers clusters that represent interpretable breakthroughs in the model's capabilities. We demonstrate the promise of these hidden phase transitions as a tool for unsupervised interpretability.
title Hidden Breakthroughs in Language Model Training
topic Machine Learning
url https://arxiv.org/abs/2506.15872