Non-asymptotic Convergence of Training Transformers for Next-token Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Ruiquan, Liang, Yingbin, Yang, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
di: Huang, Ruiquan, et al.
Pubblicazione: (2024)
di: Huang, Ruiquan, et al.
Pubblicazione: (2024)
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
di: Huang, Ruiquan, et al.
Pubblicazione: (2023)
di: Huang, Ruiquan, et al.
Pubblicazione: (2023)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
di: Huang, Ruiquan, et al.
Pubblicazione: (2026)
di: Huang, Ruiquan, et al.
Pubblicazione: (2026)
Physics in Next-token Prediction
di: An, Hongjun, et al.
Pubblicazione: (2024)
di: An, Hongjun, et al.
Pubblicazione: (2024)
Federated Online Prediction from Experts with Differential Privacy: Separations and Regret Speed-ups
di: Gao, Fengyu, et al.
Pubblicazione: (2024)
di: Gao, Fengyu, et al.
Pubblicazione: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
di: Yang, Tong, et al.
Pubblicazione: (2024)
di: Yang, Tong, et al.
Pubblicazione: (2024)
Sharp Convergence Rates for Masked Diffusion Models
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
di: Yang, Hongru, et al.
Pubblicazione: (2024)
di: Yang, Hongru, et al.
Pubblicazione: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Absorb and Converge: Provable Convergence Guarantee for Absorbing Discrete Diffusion Models
di: Liang, Yuchen, et al.
Pubblicazione: (2025)
di: Liang, Yuchen, et al.
Pubblicazione: (2025)
Provable In-Context Learning of Nonlinear Regression with Transformers
di: Li, Hongbo, et al.
Pubblicazione: (2025)
di: Li, Hongbo, et al.
Pubblicazione: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
di: Yang, Tong, et al.
Pubblicazione: (2026)
di: Yang, Tong, et al.
Pubblicazione: (2026)
Next-token pretraining implies in-context learning
di: Riechers, Paul M., et al.
Pubblicazione: (2025)
di: Riechers, Paul M., et al.
Pubblicazione: (2025)
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
di: Shen, Wei, et al.
Pubblicazione: (2024)
di: Shen, Wei, et al.
Pubblicazione: (2024)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
di: Huang, Yu, et al.
Pubblicazione: (2024)
di: Huang, Yu, et al.
Pubblicazione: (2024)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
di: Zhao, Yize, et al.
Pubblicazione: (2024)
di: Zhao, Yize, et al.
Pubblicazione: (2024)
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
di: Huang, Ruiquan, et al.
Pubblicazione: (2025)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
di: Yang, Tong, et al.
Pubblicazione: (2025)
di: Yang, Tong, et al.
Pubblicazione: (2025)
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
di: Liang, Yuchen, et al.
Pubblicazione: (2026)
Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
di: Li, Hongbo, et al.
Pubblicazione: (2025)
di: Li, Hongbo, et al.
Pubblicazione: (2025)
Transformers Provably Learn Directed Acyclic Graphs via Kernel-Guided Mutual Information
di: Cheng, Yuan, et al.
Pubblicazione: (2025)
di: Cheng, Yuan, et al.
Pubblicazione: (2025)
Tokenphormer: Structure-aware Multi-token Graph Transformer for Node Classification
di: Zhou, Zijie, et al.
Pubblicazione: (2024)
di: Zhou, Zijie, et al.
Pubblicazione: (2024)
Constraint-Rectified Training for Efficient Chain-of-Thought
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
di: Wu, Qinhang, et al.
Pubblicazione: (2026)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
Theory on Score-Mismatched Diffusion Models and Zero-Shot Conditional Samplers
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
di: Liang, Yuchen, et al.
Pubblicazione: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
di: Sander, Michael E., et al.
Pubblicazione: (2024)
di: Sander, Michael E., et al.
Pubblicazione: (2024)
Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization
di: Chen, Xuxi, et al.
Pubblicazione: (2024)
di: Chen, Xuxi, et al.
Pubblicazione: (2024)
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation
di: Pan, Pei-Chi, et al.
Pubblicazione: (2026)
di: Pan, Pei-Chi, et al.
Pubblicazione: (2026)
Non-Parametric Learning of Stochastic Differential Equations with Non-asymptotic Fast Rates of Convergence
di: Bonalli, Riccardo, et al.
Pubblicazione: (2023)
di: Bonalli, Riccardo, et al.
Pubblicazione: (2023)
Not all tokens are needed(NAT): token efficient reinforcement learning
di: Sang, Hejian, et al.
Pubblicazione: (2026)
di: Sang, Hejian, et al.
Pubblicazione: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
Next-Latent Prediction Transformers Learn Compact World Models
di: Teoh, Jayden, et al.
Pubblicazione: (2025)
di: Teoh, Jayden, et al.
Pubblicazione: (2025)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
di: Jin, Ruinan, et al.
Pubblicazione: (2026)
Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?
di: Ju, Peizhong, et al.
Pubblicazione: (2024)
di: Ju, Peizhong, et al.
Pubblicazione: (2024)
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
di: Yang, Hongru, et al.
Pubblicazione: (2023)
di: Yang, Hongru, et al.
Pubblicazione: (2023)
Flow Matching for Offline Reinforcement Learning with Discrete Actions
di: Khan, Fairoz Nower, et al.
Pubblicazione: (2026)
di: Khan, Fairoz Nower, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
di: Huang, Ruiquan, et al.
Pubblicazione: (2025) -
Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
di: Huang, Ruiquan, et al.
Pubblicazione: (2024) -
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
di: Huang, Ruiquan, et al.
Pubblicazione: (2023) -
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
di: Huang, Ruiquan, et al.
Pubblicazione: (2026) -
Physics in Next-token Prediction
di: An, Hongjun, et al.
Pubblicazione: (2024)