Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Gagnon, Leo, Elmoznino, Eric, Mittal, Sarthak, Marty, Tom, Kasetty, Tejas, Sridhar, Dhanya, Lajoie, Guillaume |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Compression Perspective on Simplicity Bias
by: Marty, Tom, et al.
Published: (2026)
by: Marty, Tom, et al.
Published: (2026)
In-context learning and Occam's razor
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Does learning the right latent variables necessarily improve in-context learning?
by: Mittal, Sarthak, et al.
Published: (2024)
by: Mittal, Sarthak, et al.
Published: (2024)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Discrete, compositional, and symbolic representations through attractor dynamics
by: Nam, Andrew, et al.
Published: (2023)
by: Nam, Andrew, et al.
Published: (2023)
Amortized In-Context Bayesian Posterior Estimation
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Multi-agent cooperation through learning-aware policy gradients
by: Meulemans, Alexander, et al.
Published: (2024)
by: Meulemans, Alexander, et al.
Published: (2024)
Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
by: Rhee, Phill Kyu
Published: (2025)
by: Rhee, Phill Kyu
Published: (2025)
In-Context Imitation Learning via Next-Token Prediction
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
Engineering Sentience
by: Demin, Konstantin, et al.
Published: (2025)
by: Demin, Konstantin, et al.
Published: (2025)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025)
by: Joshi, Shruti, et al.
Published: (2025)
A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases
by: Lombardo, Gianfranco, et al.
Published: (2026)
by: Lombardo, Gianfranco, et al.
Published: (2026)
Provable Long-Range Benefits of Next-Token Prediction
by: Cao, Xinyuan, et al.
Published: (2025)
by: Cao, Xinyuan, et al.
Published: (2025)
Next-Token Prediction and Regret Minimization
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
How Should We Meta-Learn Reinforcement Learning Algorithms?
by: Goldie, Alexander David, et al.
Published: (2025)
by: Goldie, Alexander David, et al.
Published: (2025)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
Alternatives To Next Token Prediction In Text Generation -- A Survey
by: Wyatt, Charlie, et al.
Published: (2025)
by: Wyatt, Charlie, et al.
Published: (2025)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
Creative Loss: Ambiguity, Uncertainty and Indeterminacy
by: Holberton, Tom
Published: (2024)
by: Holberton, Tom
Published: (2024)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction
by: Yin, Xuanhua, et al.
Published: (2026)
by: Yin, Xuanhua, et al.
Published: (2026)
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
by: Yang, Chun-Hao, et al.
Published: (2025)
by: Yang, Chun-Hao, et al.
Published: (2025)
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
by: He, Jianping, et al.
Published: (2025)
by: He, Jianping, et al.
Published: (2025)
Modeling Next-Token Prediction as Left-Nested Intuitionistic Implication
by: Tarau, Paul
Published: (2026)
by: Tarau, Paul
Published: (2026)
LLMs are Not Just Next Token Predictors
by: Downes, Stephen M., et al.
Published: (2024)
by: Downes, Stephen M., et al.
Published: (2024)
Agentic AI Systems Should Be Designed as Marginal Token Allocators
by: Zhu, Siqi
Published: (2026)
by: Zhu, Siqi
Published: (2026)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
by: Chen, Yuzhu, et al.
Published: (2025)
by: Chen, Yuzhu, et al.
Published: (2025)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
by: Jia, Mumin, et al.
Published: (2025)
by: Jia, Mumin, et al.
Published: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
by: An, Chenyang, et al.
Published: (2024)
by: An, Chenyang, et al.
Published: (2024)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Manifold Trajectories in Next-Token Prediction: From Replicator Dynamics to Softmax Equilibrium
by: Lee-Jenkins, Christopher R.
Published: (2025)
by: Lee-Jenkins, Christopher R.
Published: (2025)
Similar Items
-
A Compression Perspective on Simplicity Bias
by: Marty, Tom, et al.
Published: (2026) -
In-context learning and Occam's razor
by: Elmoznino, Eric, et al.
Published: (2024) -
Does learning the right latent variables necessarily improve in-context learning?
by: Mittal, Sarthak, et al.
Published: (2024) -
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026) -
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)