Controlling changes to attention logits
Fuente:
arXiv
Saved in:
| Main Authors: | Anson, Ben, Aitchison, Laurence |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025)
by: Milsom, Edward, et al.
Published: (2025)
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023)
by: Milsom, Edward, et al.
Published: (2023)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024)
by: Milsom, Edward, et al.
Published: (2024)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Using Neural Networks for Data Cleaning in Weather Datasets
by: Hanslope, Jack R. P., et al.
Published: (2024)
by: Hanslope, Jack R. P., et al.
Published: (2024)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Inverse-Free Sparse Variational Gaussian Processes
by: Cortinovis, Stefano, et al.
Published: (2026)
by: Cortinovis, Stefano, et al.
Published: (2026)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
by: Fox, David, et al.
Published: (2026)
by: Fox, David, et al.
Published: (2026)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
by: Kanai, Sekitoshi, et al.
Published: (2025)
by: Kanai, Sekitoshi, et al.
Published: (2025)
Latent attention on masked patches for flow reconstruction
by: Eze, Ben, et al.
Published: (2026)
by: Eze, Ben, et al.
Published: (2026)
Questionable practices in machine learning
by: Leech, Gavin, et al.
Published: (2024)
by: Leech, Gavin, et al.
Published: (2024)
Graph neural networks for residential location choice: connection to classical logit models
by: Cheng, Zhanhong, et al.
Published: (2025)
by: Cheng, Zhanhong, et al.
Published: (2025)
Machine learning emulation of precipitation from km-scale UK regional climate simulations using a diffusion model
by: Addison, Henry, et al.
Published: (2024)
by: Addison, Henry, et al.
Published: (2024)
Bayesian Reward Models for LLM Alignment
by: Yang, Adam X., et al.
Published: (2024)
by: Yang, Adam X., et al.
Published: (2024)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
by: Gan, Xingwei, et al.
Published: (2026)
by: Gan, Xingwei, et al.
Published: (2026)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Easy attention: A simple attention mechanism for temporal predictions with transformers
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
From independent patches to coordinated attention: Controlling information flow in vision transformers
by: Murphy, Kieran A.
Published: (2026)
by: Murphy, Kieran A.
Published: (2026)
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
Compete and Compose: Learning Independent Mechanisms for Modular World Models
by: Lei, Anson, et al.
Published: (2024)
by: Lei, Anson, et al.
Published: (2024)
Disentangling Dynamical Systems: Causal Representation Learning Meets Local Sparse Attention
by: Baumgartner, Markus W., et al.
Published: (2026)
by: Baumgartner, Markus W., et al.
Published: (2026)
Offline-to-online Reinforcement Learning for Image-based Grasping with Scarce Demonstrations
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
Approximation of relation functions and attention mechanisms
by: Altabaa, Awni, et al.
Published: (2024)
by: Altabaa, Awni, et al.
Published: (2024)
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
Beyond Spatio-Temporal Representations: Evolving Fourier Transform for Temporal Graphs
by: Bastos, Anson, et al.
Published: (2024)
by: Bastos, Anson, et al.
Published: (2024)
Myosotis: structured computation for attention like layer
by: Egorov, Evgenii, et al.
Published: (2025)
by: Egorov, Evgenii, et al.
Published: (2025)
Fast attention mechanisms: a tale of parallelism
by: Liu, Jingwen, et al.
Published: (2025)
by: Liu, Jingwen, et al.
Published: (2025)
An extension of linear self-attention for in-context learning
by: Hagiwara, Katsuyuki
Published: (2025)
by: Hagiwara, Katsuyuki
Published: (2025)
Supervised learning pays attention
by: Craig, Erin, et al.
Published: (2025)
by: Craig, Erin, et al.
Published: (2025)
Similar Items
-
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025) -
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024) -
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023) -
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024) -
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)