The Information of Large Language Model Geometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhiquan, Li, Chenghai, Huang, Weiran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929234692276224
author Tan, Zhiquan
Li, Chenghai
Huang, Weiran
author_facet Tan, Zhiquan
Li, Chenghai
Huang, Weiran
contents This paper investigates the information encoded in the embeddings of large language models (LLMs). We conduct simulations to analyze the representation entropy and discover a power law relationship with model sizes. Building upon this observation, we propose a theory based on (conditional) entropy to elucidate the scaling law phenomenon. Furthermore, we delve into the auto-regressive structure of LLMs and examine the relationship between the last token and previous context tokens using information theory and regression techniques. Specifically, we establish a theoretical connection between the information gain of new tokens and ridge regression. Additionally, we explore the effectiveness of Lasso regression in selecting meaningful tokens, which sometimes outperforms the closely related attention weights. Finally, we conduct controlled experiments, and find that information is distributed across tokens, rather than being concentrated in specific "meaningful" tokens alone.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03471
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Information of Large Language Model Geometry
Tan, Zhiquan
Li, Chenghai
Huang, Weiran
Machine Learning
Artificial Intelligence
Computation and Language
Information Theory
This paper investigates the information encoded in the embeddings of large language models (LLMs). We conduct simulations to analyze the representation entropy and discover a power law relationship with model sizes. Building upon this observation, we propose a theory based on (conditional) entropy to elucidate the scaling law phenomenon. Furthermore, we delve into the auto-regressive structure of LLMs and examine the relationship between the last token and previous context tokens using information theory and regression techniques. Specifically, we establish a theoretical connection between the information gain of new tokens and ridge regression. Additionally, we explore the effectiveness of Lasso regression in selecting meaningful tokens, which sometimes outperforms the closely related attention weights. Finally, we conduct controlled experiments, and find that information is distributed across tokens, rather than being concentrated in specific "meaningful" tokens alone.
title The Information of Large Language Model Geometry
topic Machine Learning
Artificial Intelligence
Computation and Language
Information Theory
url https://arxiv.org/abs/2402.03471