How Well Can Transformers Emulate In-context Newton's Method?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Giannou, Angeliki, Yang, Liu, Wang, Tianhao, Papailiopoulos, Dimitris, Lee, Jason D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks
von: Cayci, Semih
Veröffentlicht: (2024)
von: Cayci, Semih
Veröffentlicht: (2024)
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
Frugality in second-order optimization: floating-point approximations for Newton's method
von: Carrino, Giuseppe, et al.
Veröffentlicht: (2025)
von: Carrino, Giuseppe, et al.
Veröffentlicht: (2025)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime
von: Geng, Haoyu, et al.
Veröffentlicht: (2023)
von: Geng, Haoyu, et al.
Veröffentlicht: (2023)
Understanding Optimization in Deep Learning with Central Flows
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
von: Liu, Xin, et al.
Veröffentlicht: (2022)
von: Liu, Xin, et al.
Veröffentlicht: (2022)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
von: Yang, Tong, et al.
Veröffentlicht: (2026)
von: Yang, Tong, et al.
Veröffentlicht: (2026)
On Some Tunable Multi-fidelity Bayesian Optimization Frameworks
von: Manoj, Arjun, et al.
Veröffentlicht: (2025)
von: Manoj, Arjun, et al.
Veröffentlicht: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
How to escape sharp minima with random perturbations
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
Stability of Transformers under Layer Normalization
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
Training Infinitely Deep and Wide Transformers
von: Barboni, Raphaël, et al.
Veröffentlicht: (2026)
von: Barboni, Raphaël, et al.
Veröffentlicht: (2026)
Spatial Transformers for Radio Map Estimation
von: Viet, Pham Q., et al.
Veröffentlicht: (2024)
von: Viet, Pham Q., et al.
Veröffentlicht: (2024)
DualSchool: How Reliable are LLMs for Optimization Education?
von: Klamkin, Michael, et al.
Veröffentlicht: (2025)
von: Klamkin, Michael, et al.
Veröffentlicht: (2025)
A multilevel approach to accelerate the training of Transformers
von: Lauga, Guillaume, et al.
Veröffentlicht: (2025)
von: Lauga, Guillaume, et al.
Veröffentlicht: (2025)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
von: Feng, Mingquan, et al.
Veröffentlicht: (2024)
von: Feng, Mingquan, et al.
Veröffentlicht: (2024)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
von: Zimin, Aleksandr, et al.
Veröffentlicht: (2026)
von: Zimin, Aleksandr, et al.
Veröffentlicht: (2026)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
von: Arda, Enes, et al.
Veröffentlicht: (2026)
von: Arda, Enes, et al.
Veröffentlicht: (2026)
On the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
von: Huang, Yu, et al.
Veröffentlicht: (2025)
von: Huang, Yu, et al.
Veröffentlicht: (2025)
Adaptive Primal-Dual Method for Safe Reinforcement Learning
von: Chen, Weiqin, et al.
Veröffentlicht: (2024)
von: Chen, Weiqin, et al.
Veröffentlicht: (2024)
Reward Collapse in Aligning Large Language Models
von: Song, Ziang, et al.
Veröffentlicht: (2023)
von: Song, Ziang, et al.
Veröffentlicht: (2023)
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
von: Hu, Hanyu, et al.
Veröffentlicht: (2025)
von: Hu, Hanyu, et al.
Veröffentlicht: (2025)
Double Momentum Method for Lower-Level Constrained Bilevel Optimization
von: Shi, Wanli, et al.
Veröffentlicht: (2024)
von: Shi, Wanli, et al.
Veröffentlicht: (2024)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
von: Shen, Jucheng, et al.
Veröffentlicht: (2026)
von: Shen, Jucheng, et al.
Veröffentlicht: (2026)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
von: Xiao, Nachuan, et al.
Veröffentlicht: (2023)
Federated Instrumental Variable Analysis via Federated Generalized Method of Moments
von: Geetika, et al.
Veröffentlicht: (2025)
von: Geetika, et al.
Veröffentlicht: (2025)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
von: Ameli, Mostafa, et al.
Veröffentlicht: (2025)
von: Ameli, Mostafa, et al.
Veröffentlicht: (2025)
Unbiased Gradient Low-Rank Projection
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024) -
A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks
von: Cayci, Semih
Veröffentlicht: (2024) -
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026) -
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
von: Li, Zihao, et al.
Veröffentlicht: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)