Evolution of Concepts in Language Model Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911447186931712 |
|---|---|
| author | Ge, Xuyang Shu, Wentao Wu, Jiaxing Zhou, Yunhua He, Zhengfu Qiu, Xipeng |
| author_facet | Ge, Xuyang Shu, Wentao Wu, Jiaxing Zhou, Yunhua He, Zhengfu Qiu, Xipeng |
| contents | Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17196 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Evolution of Concepts in Language Model Pre-Training Ge, Xuyang Shu, Wentao Wu, Jiaxing Zhou, Yunhua He, Zhengfu Qiu, Xipeng Computation and Language Artificial Intelligence Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics. |
| title | Evolution of Concepts in Language Model Pre-Training |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2509.17196 |