Tracing the Representation Geometry of Language Models from Pretraining to Post-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Melody Zixuan, Agrawal, Kumar Krishna, Ghosh, Arna, Teru, Komal Kumar, Santoro, Adam, Lajoie, Guillaume, Richards, Blake A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918149224398848
author Li, Melody Zixuan
Agrawal, Kumar Krishna
Ghosh, Arna
Teru, Komal Kumar
Santoro, Adam
Lajoie, Guillaume
Richards, Blake A.
author_facet Li, Melody Zixuan
Agrawal, Kumar Krishna
Ghosh, Arna
Teru, Komal Kumar
Santoro, Adam
Lajoie, Guillaume
Richards, Blake A.
contents Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learned representations across pretraining and post-training, measuring effective rank (RankMe) and eigenspectrum decay ($α$-ReQ). With OLMo (1B-7B) and Pythia (160M-12B) models, we uncover a consistent non-monotonic sequence of three geometric phases during autoregressive pretraining. The initial "warmup" phase exhibits rapid representational collapse. This is followed by an "entropy-seeking" phase, where the manifold's dimensionality expands substantially, coinciding with peak n-gram memorization. Subsequently, a "compression-seeking" phase imposes anisotropic consolidation, selectively preserving variance along dominant eigendirections while contracting others, a transition marked with significant improvement in downstream task performance. We show these phases can emerge from a fundamental interplay of cross-entropy optimization under skewed token frequencies and representational bottlenecks ($d \ll |V|$). Post-training further transforms geometry: SFT and DPO drive "entropy-seeking" dynamics to integrate specific instructional or preferential data, improving in-distribution performance while degrading out-of-distribution robustness. Conversely, RLVR induces "compression-seeking", enhancing reward alignment but reducing generation diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tracing the Representation Geometry of Language Models from Pretraining to Post-training
Li, Melody Zixuan
Agrawal, Kumar Krishna
Ghosh, Arna
Teru, Komal Kumar
Santoro, Adam
Lajoie, Guillaume
Richards, Blake A.
Machine Learning
Artificial Intelligence
Computation and Language
Standard training metrics like loss fail to explain the emergence of complex capabilities in large language models. We take a spectral approach to investigate the geometry of learned representations across pretraining and post-training, measuring effective rank (RankMe) and eigenspectrum decay ($α$-ReQ). With OLMo (1B-7B) and Pythia (160M-12B) models, we uncover a consistent non-monotonic sequence of three geometric phases during autoregressive pretraining. The initial "warmup" phase exhibits rapid representational collapse. This is followed by an "entropy-seeking" phase, where the manifold's dimensionality expands substantially, coinciding with peak n-gram memorization. Subsequently, a "compression-seeking" phase imposes anisotropic consolidation, selectively preserving variance along dominant eigendirections while contracting others, a transition marked with significant improvement in downstream task performance. We show these phases can emerge from a fundamental interplay of cross-entropy optimization under skewed token frequencies and representational bottlenecks ($d \ll |V|$). Post-training further transforms geometry: SFT and DPO drive "entropy-seeking" dynamics to integrate specific instructional or preferential data, improving in-distribution performance while degrading out-of-distribution robustness. Conversely, RLVR induces "compression-seeking", enhancing reward alignment but reducing generation diversity.
title Tracing the Representation Geometry of Language Models from Pretraining to Post-training
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.23024