Auto-Regressive Next-Token Predictors are Universal Learners
Fuente:
arXiv
Saved in:
| Main Author: | Malach, Eran |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
On the Power of Decision Trees in Auto-Regressive Language Modeling
by: Gan, Yulu, et al.
Published: (2024)
by: Gan, Yulu, et al.
Published: (2024)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025)
by: Tsilivis, Nikolaos, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Controlling Summarization Length Through EOS Token Weighting
by: Belligoli, Zeno, et al.
Published: (2025)
by: Belligoli, Zeno, et al.
Published: (2025)
Boosting Masked ECG-Text Auto-Encoders as Discriminative Learners
by: Pham, Hung Manh, et al.
Published: (2024)
by: Pham, Hung Manh, et al.
Published: (2024)
On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
by: Alberghi, Riccardo, et al.
Published: (2025)
by: Alberghi, Riccardo, et al.
Published: (2025)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Transformers are Universal In-context Learners
by: Furuya, Takashi, et al.
Published: (2024)
by: Furuya, Takashi, et al.
Published: (2024)
Genomic Next-Token Predictors are In-Context Learners
by: Breslow, Nathan, et al.
Published: (2025)
by: Breslow, Nathan, et al.
Published: (2025)
Cubit: Token Mixer with Kernel Ridge Regression
by: Zheng, Chuanyang, et al.
Published: (2026)
by: Zheng, Chuanyang, et al.
Published: (2026)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
by: Plyler, Mitchell, et al.
Published: (2025)
by: Plyler, Mitchell, et al.
Published: (2025)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
by: Qi, Mengnan, et al.
Published: (2024)
by: Qi, Mengnan, et al.
Published: (2024)
Efficient Context Propagating Perceiver Architectures for Auto-Regressive Language Modeling
by: Mahmood, Kaleel, et al.
Published: (2024)
by: Mahmood, Kaleel, et al.
Published: (2024)
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
by: Lovelace, Justin, et al.
Published: (2026)
by: Lovelace, Justin, et al.
Published: (2026)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
SQLformer: Deep Auto-Regressive Query Graph Generation for Text-to-SQL Translation
by: Bazaga, Adrián, et al.
Published: (2023)
by: Bazaga, Adrián, et al.
Published: (2023)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
by: Trauger, Jacob, et al.
Published: (2025)
by: Trauger, Jacob, et al.
Published: (2025)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Differentially Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
by: Marconato, Emanuele, et al.
Published: (2024)
by: Marconato, Emanuele, et al.
Published: (2024)
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
by: Corielli, Francesco
Published: (2026)
by: Corielli, Francesco
Published: (2026)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
by: Wang, Duo, et al.
Published: (2024)
by: Wang, Duo, et al.
Published: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
by: Jia, Mumin, et al.
Published: (2025)
by: Jia, Mumin, et al.
Published: (2025)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Similar Items
-
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025) -
On the Power of Decision Trees in Auto-Regressive Language Modeling
by: Gan, Yulu, et al.
Published: (2024) -
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026) -
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025) -
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)