Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Blondel, Mathieu, Sander, Michael E., Vivier-Ardisson, Germain, Liu, Tianlin, Roulet, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Learning of Energy-based Models and their Partition Function
by: Sander, Michael E., et al.
Published: (2025)
by: Sander, Michael E., et al.
Published: (2025)
Differentiable Knapsack and Top-k Operators via Dynamic Programming
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
Learning with Local Search MCMC Layers
by: Vivier-Ardisson, Germain, et al.
Published: (2025)
by: Vivier-Ardisson, Germain, et al.
Published: (2025)
Regularized Large Neighborhood Search
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
by: Vivier-Ardisson, Germain, et al.
Published: (2026)
Loss Functions and Operators Generated by f-Divergences
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
The Elements of Differentiable Programming
by: Blondel, Mathieu, et al.
Published: (2024)
by: Blondel, Mathieu, et al.
Published: (2024)
CF-OPT: Counterfactual Explanations for Structured Prediction
by: Vivier-Ardisson, Germain, et al.
Published: (2024)
by: Vivier-Ardisson, Germain, et al.
Published: (2024)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
Routers in Vision Mixture of Experts: An Empirical Study
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
Stepping on the Edge: Curvature Aware Learning Rate Tuners
by: Roulet, Vincent, et al.
Published: (2024)
by: Roulet, Vincent, et al.
Published: (2024)
Next-Depth Lookahead Tree
by: Lee, Jaeho, et al.
Published: (2025)
by: Lee, Jaeho, et al.
Published: (2025)
Multi-scale Graph Autoregressive Modeling: Molecular Property Prediction via Next Token Prediction
by: Jiang, Zhuoyang, et al.
Published: (2026)
by: Jiang, Zhuoyang, et al.
Published: (2026)
Decoding-time Realignment of Language Models
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
Adaptively Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
by: Chen, Yuzhu, et al.
Published: (2025)
by: Chen, Yuzhu, et al.
Published: (2025)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Differentially Private Next-Token Prediction of Large Language Models
by: Flemings, James, et al.
Published: (2024)
by: Flemings, James, et al.
Published: (2024)
Trajeglish: Traffic Modeling as Next-Token Prediction
by: Philion, Jonah, et al.
Published: (2023)
by: Philion, Jonah, et al.
Published: (2023)
Masked Diffusion Models are Secretly Learned-Order Autoregressive Models
by: Garg, Prateek, et al.
Published: (2025)
by: Garg, Prateek, et al.
Published: (2025)
Generative Verifiers: Reward Modeling as Next-Token Prediction
by: Zhang, Lunjun, et al.
Published: (2024)
by: Zhang, Lunjun, et al.
Published: (2024)
Lookahead Drifting Model
by: Zhang, Guoqiang, et al.
Published: (2026)
by: Zhang, Guoqiang, et al.
Published: (2026)
Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
by: Zhong, Qimin, et al.
Published: (2025)
by: Zhong, Qimin, et al.
Published: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
by: Rohatgi, Dhruv, et al.
Published: (2025)
by: Rohatgi, Dhruv, et al.
Published: (2025)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
by: Wang, Siwei, et al.
Published: (2024)
by: Wang, Siwei, et al.
Published: (2024)
Per-example gradients: a new frontier for understanding and improving optimizers
by: Roulet, Vincent, et al.
Published: (2025)
by: Roulet, Vincent, et al.
Published: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases
by: Lombardo, Gianfranco, et al.
Published: (2026)
by: Lombardo, Gianfranco, et al.
Published: (2026)
Prot2Token: A Unified Framework for Protein Modeling via Next-Token Prediction
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
NoLBERT: A No Lookahead(back) Foundational Language Model
by: Kakhbod, Ali, et al.
Published: (2025)
by: Kakhbod, Ali, et al.
Published: (2025)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)
by: Jelassi, Samy, et al.
Published: (2026)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Enabling Autoregressive Models to Fill In Masked Tokens
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Similar Items
-
Joint Learning of Energy-based Models and their Partition Function
by: Sander, Michael E., et al.
Published: (2025) -
Differentiable Knapsack and Top-k Operators via Dynamic Programming
by: Vivier-Ardisson, Germain, et al.
Published: (2026) -
Learning with Local Search MCMC Layers
by: Vivier-Ardisson, Germain, et al.
Published: (2025) -
Regularized Large Neighborhood Search
by: Vivier-Ardisson, Germain, et al.
Published: (2026) -
Loss Functions and Operators Generated by f-Divergences
by: Roulet, Vincent, et al.
Published: (2025)