Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yize, Behnia, Tina, Vakilian, Vala, Thrampoulidis, Christos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916619866865664
author Zhao, Yize
Behnia, Tina
Vakilian, Vala
Thrampoulidis, Christos
author_facet Zhao, Yize
Behnia, Tina
Vakilian, Vala
Thrampoulidis, Christos
contents Next-token prediction (NTP) over large text corpora has become the go-to paradigm to train large language models. Yet, it remains unclear how NTP influences the mapping of linguistic patterns to geometric properties of the resulting model representations. We frame training of large language models as soft-label classification over sparse probabilistic label vectors, coupled with an analytical approximation that allows unrestricted generation of context embeddings. This approach links NTP training to rank-constrained, nuclear-norm regularized optimization in the logit domain, offering a framework for analyzing the geometry of word and context embeddings. In large embedding spaces, we find that NTP implicitly favors learning logits with a sparse plus low-rank structure. While the sparse component captures the co-occurrence frequency of context-word pairs, the orthogonal low-rank component, which becomes dominant as training progresses, depends solely on the sparsity pattern of the co-occurrence matrix. Consequently, when projected onto an appropriate subspace, representations of contexts that are followed by the same set of next-tokens collapse, a phenomenon we term subspace-collapse. We validate our findings on synthetic and small-scale real language datasets. Finally, we outline potential research directions aimed at deepening the understanding of NTP's influence on the learning of linguistic patterns and regularities.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15417
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
Zhao, Yize
Behnia, Tina
Vakilian, Vala
Thrampoulidis, Christos
Computation and Language
Machine Learning
Next-token prediction (NTP) over large text corpora has become the go-to paradigm to train large language models. Yet, it remains unclear how NTP influences the mapping of linguistic patterns to geometric properties of the resulting model representations. We frame training of large language models as soft-label classification over sparse probabilistic label vectors, coupled with an analytical approximation that allows unrestricted generation of context embeddings. This approach links NTP training to rank-constrained, nuclear-norm regularized optimization in the logit domain, offering a framework for analyzing the geometry of word and context embeddings. In large embedding spaces, we find that NTP implicitly favors learning logits with a sparse plus low-rank structure. While the sparse component captures the co-occurrence frequency of context-word pairs, the orthogonal low-rank component, which becomes dominant as training progresses, depends solely on the sparsity pattern of the co-occurrence matrix. Consequently, when projected onto an appropriate subspace, representations of contexts that are followed by the same set of next-tokens collapse, a phenomenon we term subspace-collapse. We validate our findings on synthetic and small-scale real language datasets. Finally, we outline potential research directions aimed at deepening the understanding of NTP's influence on the learning of linguistic patterns and regularities.
title Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2408.15417