Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rofin, Mark, Naghiyev, Jalal, Hahn, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914393603702784
author Rofin, Mark
Naghiyev, Jalal
Hahn, Michael
author_facet Rofin, Mark
Naghiyev, Jalal
Hahn, Michael
contents Trained Transformers have been shown to compute abstract features that appear redundant for predicting the immediate next token. We identify which components of the gradient signal from the next-token prediction objective give rise to this phenomenon, and we propose a method to estimate the influence of those components on the emergence of specific features. After validating our approach on toy tasks, we use it to interpret the origins of the world model in OthelloGPT and syntactic features in a small language model. Finally, we apply our framework to a pretrained LLM, showing that features with extremely high or low influence on future tokens tend to be related to formal reasoning domains such as code. Overall, our work takes a step toward understanding hidden features of Transformers through the lens of their development during training.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14087
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
Rofin, Mark
Naghiyev, Jalal
Hahn, Michael
Machine Learning
Computation and Language
Trained Transformers have been shown to compute abstract features that appear redundant for predicting the immediate next token. We identify which components of the gradient signal from the next-token prediction objective give rise to this phenomenon, and we propose a method to estimate the influence of those components on the emergence of specific features. After validating our approach on toy tasks, we use it to interpret the origins of the world model in OthelloGPT and syntactic features in a small language model. Finally, we apply our framework to a pretrained LLM, showing that features with extremely high or low influence on future tokens tend to be related to formal reasoning domains such as code. Overall, our work takes a step toward understanding hidden features of Transformers through the lens of their development during training.
title Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2603.14087