Evolution of Concepts in Language Model Pre-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Xuyang, Shu, Wentao, Wu, Jiaxing, Zhou, Yunhua, He, Zhengfu, Qiu, Xipeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911447186931712
author Ge, Xuyang
Shu, Wentao
Wu, Jiaxing
Zhou, Yunhua
He, Zhengfu
Qiu, Xipeng
author_facet Ge, Xuyang
Shu, Wentao
Wu, Jiaxing
Zhou, Yunhua
He, Zhengfu
Qiu, Xipeng
contents Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolution of Concepts in Language Model Pre-Training
Ge, Xuyang
Shu, Wentao
Wu, Jiaxing
Zhou, Yunhua
He, Zhengfu
Qiu, Xipeng
Computation and Language
Artificial Intelligence
Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics.
title Evolution of Concepts in Language Model Pre-Training
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.17196