Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gong, Zixuan, Hu, Xiaolin, Tang, Huayi, Liu, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023)
von: Malach, Eran
Veröffentlicht: (2023)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
von: Schneider, Johannes
Veröffentlicht: (2024)
von: Schneider, Johannes
Veröffentlicht: (2024)
Reasoning Bias of Next Token Prediction Training
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)
ENTP: Encoder-only Next Token Prediction
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
Efficient Context Propagating Perceiver Architectures for Auto-Regressive Language Modeling
von: Mahmood, Kaleel, et al.
Veröffentlicht: (2024)
von: Mahmood, Kaleel, et al.
Veröffentlicht: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
Enhancing LLM Safety via Constrained Direct Preference Optimization
von: Liu, Zixuan, et al.
Veröffentlicht: (2024)
von: Liu, Zixuan, et al.
Veröffentlicht: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
von: Qi, Mengnan, et al.
Veröffentlicht: (2024)
von: Qi, Mengnan, et al.
Veröffentlicht: (2024)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
SQLformer: Deep Auto-Regressive Query Graph Generation for Text-to-SQL Translation
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
von: Bazaga, Adrián, et al.
Veröffentlicht: (2023)
Cubit: Token Mixer with Kernel Ridge Regression
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2026)
A Law of Next-Token Prediction in Large Language Models
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
Differentially Private Next-Token Prediction of Large Language Models
von: Flemings, James, et al.
Veröffentlicht: (2024)
von: Flemings, James, et al.
Veröffentlicht: (2024)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
SASA: Semantic-Aware Contrastive Learning Framework with Separated Attention for Triple Classification
von: Xiaodan, Xu, et al.
Veröffentlicht: (2026)
von: Xiaodan, Xu, et al.
Veröffentlicht: (2026)
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
von: Liu, Shang, et al.
Veröffentlicht: (2024)
von: Liu, Shang, et al.
Veröffentlicht: (2024)
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
von: Su, Ye, et al.
Veröffentlicht: (2026)
von: Su, Ye, et al.
Veröffentlicht: (2026)
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
von: Jia, Mumin, et al.
Veröffentlicht: (2025)
von: Jia, Mumin, et al.
Veröffentlicht: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
von: Liu, Hanbing, et al.
Veröffentlicht: (2025)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
von: Chen, Liang, et al.
Veröffentlicht: (2024)
von: Chen, Liang, et al.
Veröffentlicht: (2024)
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming
von: Corielli, Francesco
Veröffentlicht: (2026)
von: Corielli, Francesco
Veröffentlicht: (2026)
Information-Theoretic Generalization Bounds for Transductive Learning and its Applications
von: Tang, Huayi, et al.
Veröffentlicht: (2023)
von: Tang, Huayi, et al.
Veröffentlicht: (2023)
TokenShapley: Token Level Context Attribution with Shapley Value
von: Xiao, Yingtai, et al.
Veröffentlicht: (2025)
von: Xiao, Yingtai, et al.
Veröffentlicht: (2025)
Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
von: Plyler, Mitchell, et al.
Veröffentlicht: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
von: Rofin, Mark, et al.
Veröffentlicht: (2026)
AutoBencher: Towards Declarative Benchmark Construction
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2024)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2024)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
von: Jo, Dongwon, et al.
Veröffentlicht: (2026)
Nectar: Neural Estimation of Cached-Token Attention via Regression
von: Monteiro, João, et al.
Veröffentlicht: (2026)
von: Monteiro, João, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023) -
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
von: Liu, Yuhang, et al.
Veröffentlicht: (2025) -
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025) -
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
von: Schneider, Johannes
Veröffentlicht: (2024) -
Reasoning Bias of Next Token Prediction Training
von: Lin, Pengxiao, et al.
Veröffentlicht: (2025)