I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yuhang, Gong, Dong, Cai, Yichao, Gao, Erdun, Zhang, Zhen, Huang, Biwei, Gong, Mingming, Hengel, Anton van den, Shi, Javen Qinfeng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917300094894080
author Liu, Yuhang
Gong, Dong
Cai, Yichao
Gao, Erdun
Zhang, Zhen
Huang, Biwei
Gong, Mingming
Hengel, Anton van den
Shi, Javen Qinfeng
author_facet Liu, Yuhang
Gong, Dong
Cai, Yichao
Gao, Erdun
Zhang, Zhen
Huang, Biwei
Gong, Mingming
Hengel, Anton van den
Shi, Javen Qinfeng
contents Recent empirical evidence shows that LLM representations encode human-interpretable concepts. Nevertheless, the mechanisms by which these representations emerge remain largely unexplored. To shed further light on this, we introduce a novel generative model that generates tokens on the basis of such concepts formulated as latent discrete variables. Under mild conditions, even when the mapping from the latent space to the observed space is non-invertible, we establish rigorous identifiability result: the representations learned by LLMs through next-token prediction can be approximately modeled as the logarithm of the posterior probabilities of these latent discrete concepts given input context, up to an linear transformation. This theoretical finding: 1) provides evidence that LLMs capture essential underlying generative factors, 2) offers a unified and principled perspective for understanding the linear representation hypothesis, and 3) motivates a theoretically grounded approach for evaluating sparse autoencoders. Empirically, we validate our theoretical results through evaluations on both simulation data and the Pythia, Llama, and DeepSeek model families.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08980
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
Liu, Yuhang
Gong, Dong
Cai, Yichao
Gao, Erdun
Zhang, Zhen
Huang, Biwei
Gong, Mingming
Hengel, Anton van den
Shi, Javen Qinfeng
Machine Learning
Computation and Language
Recent empirical evidence shows that LLM representations encode human-interpretable concepts. Nevertheless, the mechanisms by which these representations emerge remain largely unexplored. To shed further light on this, we introduce a novel generative model that generates tokens on the basis of such concepts formulated as latent discrete variables. Under mild conditions, even when the mapping from the latent space to the observed space is non-invertible, we establish rigorous identifiability result: the representations learned by LLMs through next-token prediction can be approximately modeled as the logarithm of the posterior probabilities of these latent discrete concepts given input context, up to an linear transformation. This theoretical finding: 1) provides evidence that LLMs capture essential underlying generative factors, 2) offers a unified and principled perspective for understanding the linear representation hypothesis, and 3) motivates a theoretically grounded approach for evaluating sparse autoencoders. Empirically, we validate our theoretical results through evaluations on both simulation data and the Pythia, Llama, and DeepSeek model families.
title I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.08980