How Reinforcement Learning After Next-Token Prediction Facilitates Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tsilivis, Nikolaos, Malach, Eran, Ullrich, Karen, Kempe, Julia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023)
by: Malach, Eran
Published: (2023)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
by: Su, Jingtong, et al.
Published: (2023)
by: Su, Jingtong, et al.
Published: (2023)
On the Geometry of Regularization in Adversarial Training: High-Dimensional Asymptotics and Generalization Bounds
by: Vilucchio, Matteo, et al.
Published: (2024)
by: Vilucchio, Matteo, et al.
Published: (2024)
Attacking Bayes: On the Adversarial Robustness of Bayesian Neural Networks
by: Feng, Yunzhen, et al.
Published: (2024)
by: Feng, Yunzhen, et al.
Published: (2024)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Learning to Think from Multiple Thinkers
by: Joshi, Nirmit, et al.
Published: (2026)
by: Joshi, Nirmit, et al.
Published: (2026)
LLM Priors for ERM over Programs
by: Singhal, Shivam, et al.
Published: (2025)
by: Singhal, Shivam, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Soft Tokens, Hard Truths
by: Butt, Natasha, et al.
Published: (2025)
by: Butt, Natasha, et al.
Published: (2025)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
by: Gagnon, Leo, et al.
Published: (2025)
by: Gagnon, Leo, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Emergent properties with repeated examples
by: Charton, François, et al.
Published: (2024)
by: Charton, François, et al.
Published: (2024)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Trajeglish: Traffic Modeling as Next-Token Prediction
by: Philion, Jonah, et al.
Published: (2023)
by: Philion, Jonah, et al.
Published: (2023)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
by: Trauger, Jacob, et al.
Published: (2025)
by: Trauger, Jacob, et al.
Published: (2025)
Generative Verifiers: Reward Modeling as Next-Token Prediction
by: Zhang, Lunjun, et al.
Published: (2024)
by: Zhang, Lunjun, et al.
Published: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
How Transformers Learn to Plan via Multi-Token Prediction
by: Huang, Jianhao, et al.
Published: (2026)
by: Huang, Jianhao, et al.
Published: (2026)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Transfer Learning Study of Motion Transformer-based Trajectory Predictions
by: Ullrich, Lars, et al.
Published: (2024)
by: Ullrich, Lars, et al.
Published: (2024)
A Comprehensive Machine Learning Framework for Micromobility Demand Prediction
by: Porat, Omri, et al.
Published: (2025)
by: Porat, Omri, et al.
Published: (2025)
Contextual Intelligence The Next Leap for Reinforcement Learning
by: Biedenkapp, André
Published: (2026)
by: Biedenkapp, André
Published: (2026)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
ENTP: Encoder-only Next Token Prediction
by: Ewer, Ethan, et al.
Published: (2024)
by: Ewer, Ethan, et al.
Published: (2024)
Similar Items
-
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023) -
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024) -
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024) -
The Price of Implicit Bias in Adversarially Robust Generalization
by: Tsilivis, Nikolaos, et al.
Published: (2024) -
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
by: Su, Jingtong, et al.
Published: (2023)