Next-token pretraining implies in-context learning
Fuente:
arXiv
Saved in:
| Main Authors: | Riechers, Paul M., Bigelow, Henry R., Alt, Eric A., Shai, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rank-1 LoRAs Encode Interpretable Reasoning Signals
by: Ward, Jake, et al.
Published: (2025)
by: Ward, Jake, et al.
Published: (2025)
Physics in Next-token Prediction
by: An, Hongjun, et al.
Published: (2024)
by: An, Hongjun, et al.
Published: (2024)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Transformers learn factored representations
by: Shai, Adam, et al.
Published: (2026)
by: Shai, Adam, et al.
Published: (2026)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
In-context learning and Occam's razor
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
by: Marconato, Emanuele, et al.
Published: (2024)
by: Marconato, Emanuele, et al.
Published: (2024)
Synthetic continued pretraining
by: Yang, Zitong, et al.
Published: (2024)
by: Yang, Zitong, et al.
Published: (2024)
Does learning the right latent variables necessarily improve in-context learning?
by: Mittal, Sarthak, et al.
Published: (2024)
by: Mittal, Sarthak, et al.
Published: (2024)
Forking Paths in Neural Text Generation
by: Bigelow, Eric, et al.
Published: (2024)
by: Bigelow, Eric, et al.
Published: (2024)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
Boosting deep Reinforcement Learning using pretraining with Logical Options
by: Ye, Zihan, et al.
Published: (2026)
by: Ye, Zihan, et al.
Published: (2026)
Analyzing limits for in-context learning
by: Naim, Omar, et al.
Published: (2025)
by: Naim, Omar, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
A deep learning and machine learning approach to predict neonatal death in the context of São Paulo
by: Raihan, Mohon, et al.
Published: (2025)
by: Raihan, Mohon, et al.
Published: (2025)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data
by: Wu, Weichang, et al.
Published: (2025)
by: Wu, Weichang, et al.
Published: (2025)
Learning to Score
by: Kriger, Yogev, et al.
Published: (2025)
by: Kriger, Yogev, et al.
Published: (2025)
Neural networks leverage nominally quantum and post-quantum representations
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Constrained belief updates explain geometric structures in transformer representations
by: Piotrowski, Mateusz, et al.
Published: (2025)
by: Piotrowski, Mateusz, et al.
Published: (2025)
State- and context-dependent robotic manipulation and grasping via uncertainty-aware imitation learning
by: Winter, Tim R., et al.
Published: (2024)
by: Winter, Tim R., et al.
Published: (2024)
Bio2Token: All-atom tokenization of any biomolecular structure with Mamba
by: Liu, Andrew, et al.
Published: (2024)
by: Liu, Andrew, et al.
Published: (2024)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
by: Sokar, Ghada, et al.
Published: (2024)
by: Sokar, Ghada, et al.
Published: (2024)
TSFM in-context learning for time-series classification of bearing-health status
by: Tokic, Michel, et al.
Published: (2025)
by: Tokic, Michel, et al.
Published: (2025)
Does Flatness imply Generalization for Logistic Loss in Univariate Two-Layer ReLU Network?
by: Qiao, Dan, et al.
Published: (2025)
by: Qiao, Dan, et al.
Published: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
Self-Improving Pretraining: using post-trained models to pretrain better models
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
In-context Learning of Evolving Data Streams with Tabular Foundational Models
by: Lourenço, Afonso, et al.
Published: (2025)
by: Lourenço, Afonso, et al.
Published: (2025)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
Research Program: Theory of Learning in Dynamical Systems
by: Hazan, Elad, et al.
Published: (2025)
by: Hazan, Elad, et al.
Published: (2025)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
by: Liu, Toni J. B., et al.
Published: (2024)
by: Liu, Toni J. B., et al.
Published: (2024)
Similar Items
-
Rank-1 LoRAs Encode Interpretable Reasoning Signals
by: Ward, Jake, et al.
Published: (2025) -
Physics in Next-token Prediction
by: An, Hongjun, et al.
Published: (2024) -
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026) -
Transformers learn factored representations
by: Shai, Adam, et al.
Published: (2026) -
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)