Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913046645964800 |
|---|---|
| author | Kishino, Ryo Takase, Yusuke Oyama, Momose Yamagiwa, Hiroaki Shimodaira, Hidetoshi |
| author_facet | Kishino, Ryo Takase, Yusuke Oyama, Momose Yamagiwa, Hiroaki Shimodaira, Hidetoshi |
| contents | Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterogeneous settings. We extend this framework to training checkpoints and intermediate layers, and establish a consistent scale for KL divergence across pretraining, model size, random seeds, quantization, fine-tuning, and layers. Analysis of Pythia pretraining trajectories further shows that changes in log-likelihood space, as measured by the scaling behavior of KL divergence, are much smaller than in weight space, resulting in subdiffusive learning trajectories and early stabilization of language-model behavior despite weight drift. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_15353 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings Kishino, Ryo Takase, Yusuke Oyama, Momose Yamagiwa, Hiroaki Shimodaira, Hidetoshi Computation and Language Log-likelihood vectors define a common space for comparing language models as probability distributions, enabling unified comparisons across heterogeneous settings. We extend this framework to training checkpoints and intermediate layers, and establish a consistent scale for KL divergence across pretraining, model size, random seeds, quantization, fine-tuning, and layers. Analysis of Pythia pretraining trajectories further shows that changes in log-likelihood space, as measured by the scaling behavior of KL divergence, are much smaller than in weight space, resulting in subdiffusive learning trajectories and early stabilization of language-model behavior despite weight drift. |
| title | Establishing a Scale for Kullback-Leibler Divergence in Language Models Across Various Settings |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.15353 |