Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Mingsong, Li, Dongyang, Kulick, Charles, Tang, Sui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Explicitly Conditioned Sparsifying Transforms
von: Pătraşcu, Andrei, et al.
Veröffentlicht: (2024)
von: Pătraşcu, Andrei, et al.
Veröffentlicht: (2024)
Towards Quantifying the Preconditioning Effect of Adam
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2024)
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2024)
Geometric Data Valuation via Leverage Scores
von: Mendoza-Smith, Rodrigo
Veröffentlicht: (2025)
von: Mendoza-Smith, Rodrigo
Veröffentlicht: (2025)
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
von: Yang, Shenghao, et al.
Veröffentlicht: (2026)
von: Yang, Shenghao, et al.
Veröffentlicht: (2026)
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
von: Shumaylov, Zakhar, et al.
Veröffentlicht: (2026)
von: Shumaylov, Zakhar, et al.
Veröffentlicht: (2026)
Min-Max Optimisation for Nonconvex-Nonconcave Functions Using a Random Zeroth-Order Extragradient Algorithm
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
Scaling physics-informed hard constraints with mixture-of-experts
von: Chalapathi, Nithin, et al.
Veröffentlicht: (2024)
von: Chalapathi, Nithin, et al.
Veröffentlicht: (2024)
Parametrizing Convex Sets Using Sublinear Neural Networks
von: Martinet, Eloi
Veröffentlicht: (2026)
von: Martinet, Eloi
Veröffentlicht: (2026)
Curvature-Aware Optimization for High-Accuracy Physics-Informed Neural Networks
von: Jnini, Anas, et al.
Veröffentlicht: (2026)
von: Jnini, Anas, et al.
Veröffentlicht: (2026)
Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
A Single-Loop Gradient Descent and Perturbed Ascent Algorithm for Nonconvex Functional Constrained Optimization
von: Lu, Songtao
Veröffentlicht: (2022)
von: Lu, Songtao
Veröffentlicht: (2022)
Examining Policy Entropy of Reinforcement Learning Agents for Personalization Tasks
von: Dereventsov, Anton, et al.
Veröffentlicht: (2022)
von: Dereventsov, Anton, et al.
Veröffentlicht: (2022)
A second-order method landing on the Stiefel manifold via Newton$\unicode{x2013}$Schulz iteration
von: Xiong, Xinhui, et al.
Veröffentlicht: (2026)
von: Xiong, Xinhui, et al.
Veröffentlicht: (2026)
Maximum Principle of Optimal Probability Density Control
von: Gaby, Nathan, et al.
Veröffentlicht: (2025)
von: Gaby, Nathan, et al.
Veröffentlicht: (2025)
On the Convergence and Size Transferability of Continuous-depth Graph Neural Networks
von: Yan, Mingsong, et al.
Veröffentlicht: (2025)
von: Yan, Mingsong, et al.
Veröffentlicht: (2025)
Universal Approximation of Nonlinear Operators and Their Derivatives
von: de Feo, Filippo
Veröffentlicht: (2026)
von: de Feo, Filippo
Veröffentlicht: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
von: Fang, Yuchen, et al.
Veröffentlicht: (2026)
Data-driven Learning of Interaction Laws in Multispecies Particle Systems with Gaussian Processes: Convergence Theory and Applications
von: Feng, Jinchao, et al.
Veröffentlicht: (2025)
von: Feng, Jinchao, et al.
Veröffentlicht: (2025)
Iterative Refinement for $\ell_p$-norm Regression
von: Adil, Deeksha, et al.
Veröffentlicht: (2019)
von: Adil, Deeksha, et al.
Veröffentlicht: (2019)
Higher Order Reduced Rank Regression
von: Greenberg, Leia, et al.
Veröffentlicht: (2025)
von: Greenberg, Leia, et al.
Veröffentlicht: (2025)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
von: Dereziński, Michał, et al.
Veröffentlicht: (2026)
von: Dereziński, Michał, et al.
Veröffentlicht: (2026)
Kernel-based potential mean-field games with unbiased random Fourier $U$-statistics
von: Nakano, Yumiharu
Veröffentlicht: (2026)
von: Nakano, Yumiharu
Veröffentlicht: (2026)
Error Feedback Can Accurately Compress Preconditioners
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Minimisation of Submodular Functions Using Gaussian Zeroth-Order Random Oracles
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
Solving Dense Linear Systems Faster Than via Preconditioning
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
von: Southworth, Ben S., et al.
Veröffentlicht: (2026)
ANaGRAM: A Natural Gradient Relative to Adapted Model for efficient PINNs learning
von: Schwencke, Nilo, et al.
Veröffentlicht: (2024)
von: Schwencke, Nilo, et al.
Veröffentlicht: (2024)
Faster Linear Systems and Matrix Norm Approximation via Multi-level Sketched Preconditioning
von: Dereziński, Michał, et al.
Veröffentlicht: (2024)
von: Dereziński, Michał, et al.
Veröffentlicht: (2024)
A distributed semismooth Newton based augmented Lagrangian method for distributed optimization
von: Ma, Qihao, et al.
Veröffentlicht: (2026)
von: Ma, Qihao, et al.
Veröffentlicht: (2026)
AutoBalance: An Automatic Balancing Framework for Training Physics-Informed Neural Networks
von: An, Kang, et al.
Veröffentlicht: (2025)
von: An, Kang, et al.
Veröffentlicht: (2025)
sparseGeoHOPCA: A Geometric Solution to Sparse Higher-Order PCA Without Covariance Estimation
von: Xu, Renjie, et al.
Veröffentlicht: (2025)
von: Xu, Renjie, et al.
Veröffentlicht: (2025)
Scalable Approximate Optimal Diagonal Preconditioning
von: Gao, Wenzhi, et al.
Veröffentlicht: (2023)
von: Gao, Wenzhi, et al.
Veröffentlicht: (2023)
Learning epidemic trajectories through Kernel Operator Learning: from modelling to optimal control
von: Ziarelli, Giovanni, et al.
Veröffentlicht: (2024)
von: Ziarelli, Giovanni, et al.
Veröffentlicht: (2024)
Properties of Fixed Points of Generalised Extra Gradient Methods Applied to Min-Max Problems
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
von: Farzin, Amir Ali, et al.
Veröffentlicht: (2025)
Introduction to optimization methods for training SciML models
von: Kopaničáková, Alena, et al.
Veröffentlicht: (2026)
von: Kopaničáková, Alena, et al.
Veröffentlicht: (2026)
Practical Topics in Optimization
von: Lu, Jun
Veröffentlicht: (2025)
von: Lu, Jun
Veröffentlicht: (2025)
Designing MacPherson Suspension Architectures using Bayesian Optimization
von: Thomas, Sinnu Susan, et al.
Veröffentlicht: (2022)
von: Thomas, Sinnu Susan, et al.
Veröffentlicht: (2022)
Preconditioning transformations of adjoint systems for evolution equations
von: Tran, Brian K., et al.
Veröffentlicht: (2025)
von: Tran, Brian K., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Explicitly Conditioned Sparsifying Transforms
von: Pătraşcu, Andrei, et al.
Veröffentlicht: (2024) -
Towards Quantifying the Preconditioning Effect of Adam
von: Das, Rudrajit, et al.
Veröffentlicht: (2024) -
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2024) -
Geometric Data Valuation via Leverage Scores
von: Mendoza-Smith, Rodrigo
Veröffentlicht: (2025) -
PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network Training
von: Yang, Shenghao, et al.
Veröffentlicht: (2026)