Function-Space Learning Rates
Fuente:
arXiv
Saved in:
| Main Authors: | Milsom, Edward, Anson, Ben, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023)
by: Milsom, Edward, et al.
Published: (2023)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024)
by: Milsom, Edward, et al.
Published: (2024)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Using Neural Networks for Data Cleaning in Weather Datasets
by: Hanslope, Jack R. P., et al.
Published: (2024)
by: Hanslope, Jack R. P., et al.
Published: (2024)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Inverse-Free Sparse Variational Gaussian Processes
by: Cortinovis, Stefano, et al.
Published: (2026)
by: Cortinovis, Stefano, et al.
Published: (2026)
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
by: Fox, David, et al.
Published: (2026)
by: Fox, David, et al.
Published: (2026)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Machine learning emulation of precipitation from km-scale UK regional climate simulations using a diffusion model
by: Addison, Henry, et al.
Published: (2024)
by: Addison, Henry, et al.
Published: (2024)
Bayesian Reward Models for LLM Alignment
by: Yang, Adam X., et al.
Published: (2024)
by: Yang, Adam X., et al.
Published: (2024)
Compete and Compose: Learning Independent Mechanisms for Modular World Models
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Neural Feature Learning in Function Space
by: Xu, Xiangxiang, et al.
Published: (2023)
by: Xu, Xiangxiang, et al.
Published: (2023)
Offline-to-online Reinforcement Learning for Image-based Grasping with Scarce Demonstrations
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
Disentangling Dynamical Systems: Causal Representation Learning Meets Local Sparse Attention
by: Baumgartner, Markus W., et al.
Published: (2026)
by: Baumgartner, Markus W., et al.
Published: (2026)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Function Spaces Without Kernels: Learning Compact Hilbert Space Representations
by: Low, Su Ann, et al.
Published: (2025)
by: Low, Su Ann, et al.
Published: (2025)
Leveraging Function Space Aggregation for Federated Learning at Scale
by: Dhawan, Nikita, et al.
Published: (2023)
by: Dhawan, Nikita, et al.
Published: (2023)
Learning the Target Network in Function Space
by: Asadi, Kavosh, et al.
Published: (2024)
by: Asadi, Kavosh, et al.
Published: (2024)
Learning Orthonormal Bases for Function Spaces
by: Kamkari, Hamidreza, et al.
Published: (2026)
by: Kamkari, Hamidreza, et al.
Published: (2026)
Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules
by: Li, Binghui, et al.
Published: (2025)
by: Li, Binghui, et al.
Published: (2025)
A Closed-Form Upper Bound for Admissible Learning-Rate Steps in Belief-Space Dynamics
by: Li, Zixi, et al.
Published: (2026)
by: Li, Zixi, et al.
Published: (2026)
Learning in Function Spaces: An Unified Functional Analytic View of Supervised and Unsupervised Learning
by: Lakshmanan, K.
Published: (2026)
by: Lakshmanan, K.
Published: (2026)
Approximation Rates of Shallow Neural Networks: Barron Spaces, Activation Functions and Optimality Analysis
by: Lu, Jian, et al.
Published: (2025)
by: Lu, Jian, et al.
Published: (2025)
GRIFDIR: Graph Resolution-Invariant FEM Diffusion Models in Function Spaces over Irregular Domains
by: Rowbottom, James, et al.
Published: (2026)
by: Rowbottom, James, et al.
Published: (2026)
Function Encoders: A Principled Approach to Transfer Learning in Hilbert Spaces
by: Ingebrand, Tyler, et al.
Published: (2025)
by: Ingebrand, Tyler, et al.
Published: (2025)
Exploring Graph Mamba: A Comprehensive Survey on State-Space Models for Graph Learning
by: Atitallah, Safa Ben, et al.
Published: (2024)
by: Atitallah, Safa Ben, et al.
Published: (2024)
Neural Operator: Learning Maps Between Function Spaces
by: Kovachki, Nikola, et al.
Published: (2021)
by: Kovachki, Nikola, et al.
Published: (2021)
Learning and Blending Robot Hugging Behaviors in Time and Space
by: Drolet, Michael, et al.
Published: (2022)
by: Drolet, Michael, et al.
Published: (2022)
Similar Items
-
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024) -
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023) -
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024) -
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025) -
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)