Mechanics of Next Token Prediction with Self-Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yingcong, Huang, Yixiao, Ildiz, M. Emrullah, Rawat, Ankit Singh, Oymak, Samet |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
When and How Unlabeled Data Provably Improve In-Context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)
di: Li, Yingcong, et al.
Pubblicazione: (2025)
TimePFN: Effective Multivariate Time Series Forecasting with Synthetic Data
di: Taga, Ege Onur, et al.
Pubblicazione: (2025)
di: Taga, Ege Onur, et al.
Pubblicazione: (2025)
Retrieval Augmented Time Series Forecasting
di: Tire, Kutay, et al.
Pubblicazione: (2024)
di: Tire, Kutay, et al.
Pubblicazione: (2024)
Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
di: Ildiz, Muhammed Emrullah, et al.
Pubblicazione: (2026)
di: Ildiz, Muhammed Emrullah, et al.
Pubblicazione: (2026)
Continuous Chain of Thought Enables Parallel Exploration and Reasoning
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
di: Gozeten, Halil Alperen, et al.
Pubblicazione: (2025)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
di: Huang, Xinmeng, et al.
Pubblicazione: (2024)
Kinematic Tokenization: Optimization-Based Continuous-Time Tokens for Learnable Decision Policies in Noisy Time Series
di: Kearney, Griffin
Pubblicazione: (2026)
di: Kearney, Griffin
Pubblicazione: (2026)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
di: Gurses, Selcuk, et al.
Pubblicazione: (2025)
di: Gurses, Selcuk, et al.
Pubblicazione: (2025)
Generative AI and Process Systems Engineering: The Next Frontier
di: Decardi-Nelson, Benjamin, et al.
Pubblicazione: (2024)
di: Decardi-Nelson, Benjamin, et al.
Pubblicazione: (2024)
Can Transformers Learn Optimal Filtering for Unknown Systems?
di: Balim, Haldun, et al.
Pubblicazione: (2023)
di: Balim, Haldun, et al.
Pubblicazione: (2023)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
di: Jia, Mumin, et al.
Pubblicazione: (2025)
di: Jia, Mumin, et al.
Pubblicazione: (2025)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
di: Fernando, Heshan, et al.
Pubblicazione: (2024)
di: Fernando, Heshan, et al.
Pubblicazione: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
di: Tsiourvas, Asterios, et al.
Pubblicazione: (2025)
di: Tsiourvas, Asterios, et al.
Pubblicazione: (2025)
Reinforcement Learning from Human Feedback with Active Queries
di: Ji, Kaixuan, et al.
Pubblicazione: (2024)
di: Ji, Kaixuan, et al.
Pubblicazione: (2024)
Reward Collapse in Aligning Large Language Models
di: Song, Ziang, et al.
Pubblicazione: (2023)
di: Song, Ziang, et al.
Pubblicazione: (2023)
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
di: Gautam, Tanmay, et al.
Pubblicazione: (2024)
di: Gautam, Tanmay, et al.
Pubblicazione: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
di: Chen, Siyu, et al.
Pubblicazione: (2024)
di: Chen, Siyu, et al.
Pubblicazione: (2024)
Variational Learning is Effective for Large Deep Networks
di: Shen, Yuesong, et al.
Pubblicazione: (2024)
di: Shen, Yuesong, et al.
Pubblicazione: (2024)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
di: Pan, Rui, et al.
Pubblicazione: (2024)
di: Pan, Rui, et al.
Pubblicazione: (2024)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
di: Tao, Hongyi, et al.
Pubblicazione: (2026)
di: Tao, Hongyi, et al.
Pubblicazione: (2026)
Leveraging Large Language Models for Solving Rare MIP Challenges
di: Wang, Teng, et al.
Pubblicazione: (2024)
di: Wang, Teng, et al.
Pubblicazione: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
di: Huang, Yu, et al.
Pubblicazione: (2025)
di: Huang, Yu, et al.
Pubblicazione: (2025)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
di: Sheen, Heejune, et al.
Pubblicazione: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
di: Kharrat, Salma, et al.
Pubblicazione: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
di: Guo, Xingang, et al.
Pubblicazione: (2024)
di: Guo, Xingang, et al.
Pubblicazione: (2024)
Identification and Adaptive Control of Markov Jump Systems: Sample Complexity and Regret Bounds
di: Sattar, Yahya, et al.
Pubblicazione: (2021)
di: Sattar, Yahya, et al.
Pubblicazione: (2021)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
di: Huang, Yu, et al.
Pubblicazione: (2026)
di: Huang, Yu, et al.
Pubblicazione: (2026)
Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime
di: Geng, Haoyu, et al.
Pubblicazione: (2023)
di: Geng, Haoyu, et al.
Pubblicazione: (2023)
Differentiable Nonlinear Model Predictive Control
di: Frey, Jonathan, et al.
Pubblicazione: (2025)
di: Frey, Jonathan, et al.
Pubblicazione: (2025)
Understanding Fixed Predictions via Confined Regions
di: Lawless, Connor, et al.
Pubblicazione: (2025)
di: Lawless, Connor, et al.
Pubblicazione: (2025)
Global Convergence of Multiplicative Updates for the Matrix Mechanism: A Collaborative Proof with Gemini 3
di: Rush, Keith
Pubblicazione: (2026)
di: Rush, Keith
Pubblicazione: (2026)
Predictive and Prescriptive AI toward Optimizing Wildfire Suppression
di: Boussioux, Leonard, et al.
Pubblicazione: (2026)
di: Boussioux, Leonard, et al.
Pubblicazione: (2026)
Applications of 0-1 Neural Networks in Prescription and Prediction
di: Patil, Vrishabh, et al.
Pubblicazione: (2024)
di: Patil, Vrishabh, et al.
Pubblicazione: (2024)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
di: Chen, Zixi, et al.
Pubblicazione: (2025)
di: Chen, Zixi, et al.
Pubblicazione: (2025)
Conformal Prediction-Driven Adaptive Sampling for Digital Water Twins
di: Homaei, Mohammadhossein, et al.
Pubblicazione: (2025)
di: Homaei, Mohammadhossein, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024) -
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
di: Li, Yingcong, et al.
Pubblicazione: (2024) -
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025) -
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023) -
When and How Unlabeled Data Provably Improve In-Context Learning
di: Li, Yingcong, et al.
Pubblicazione: (2025)