What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Ormaniec, Weronika, Dangel, Felix, Singh, Sidak Pal |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Curvature Matrices Should Be Democratized via Linear Operators
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Fusion of Graph Neural Networks via Optimal Transport
by: Ormaniec, Weronika, et al.
Published: (2025)
by: Ormaniec, Weronika, et al.
Published: (2025)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Standardizing Structural Causal Models
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
by: Dangel, Felix
Published: (2023)
by: Dangel, Felix
Published: (2023)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025)
by: Theus, Alexander, et al.
Published: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Lowering PyTorch's Memory Consumption for Selective Differentiation
by: Bhatia, Samarth, et al.
Published: (2024)
by: Bhatia, Samarth, et al.
Published: (2024)
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
by: da Silva, Marvin F., et al.
Published: (2025)
by: da Silva, Marvin F., et al.
Published: (2025)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023)
by: Wenger, Jonathan, et al.
Published: (2023)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
by: Liao, Disen, et al.
Published: (2026)
by: Liao, Disen, et al.
Published: (2026)
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
by: Dangel, Felix, et al.
Published: (2024)
by: Dangel, Felix, et al.
Published: (2024)
Kronecker-factored Approximate Curvature (KFAC) From Scratch
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Collapsing Taylor Mode Automatic Differentiation
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Sketching Low-Rank Plus Diagonal Matrices
by: Fernandez, Andres, et al.
Published: (2025)
by: Fernandez, Andres, et al.
Published: (2025)
Transition Constrained Bayesian Optimization via Markov Decision Processes
by: Folch, Jose Pablo, et al.
Published: (2024)
by: Folch, Jose Pablo, et al.
Published: (2024)
Generalizing the Geometry of Model Merging Through Frechet Averages
by: da Silva, Marvin F., et al.
Published: (2026)
by: da Silva, Marvin F., et al.
Published: (2026)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
by: Milligan, Alan, et al.
Published: (2026)
by: Milligan, Alan, et al.
Published: (2026)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Theoretical Insights into Line Graph Transformation on Graph Learning
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Hessian Spectral Analysis at Foundation Model Scale
by: Granziol, Diego, et al.
Published: (2026)
by: Granziol, Diego, et al.
Published: (2026)
What Does Normal Even Mean? Evaluating Benign Traffic in Intrusion Detection Datasets
by: Wilkinson, Meghan, et al.
Published: (2025)
by: Wilkinson, Meghan, et al.
Published: (2025)
Convergence and clustering analysis for Mean Shift with radially symmetric, positive definite kernels
by: Pal, Susovan
Published: (2025)
by: Pal, Susovan
Published: (2025)
Why Transformers Need Adam: A Hessian Perspective
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
A Boolean Function-Theoretic Framework for Expressivity in GNNs with Applications to Fair Graph Mining
by: Pal, Manjish
Published: (2026)
by: Pal, Manjish
Published: (2026)
Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
by: Dixit, Shrey, et al.
Published: (2025)
by: Dixit, Shrey, et al.
Published: (2025)
Spectral-factorized Positive-definite Curvature Learning for NN Training
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Similar Items
-
Position: Curvature Matrices Should Be Democratized via Linear Operators
by: Dangel, Felix, et al.
Published: (2025) -
Fusion of Graph Neural Networks via Optimal Transport
by: Ormaniec, Weronika, et al.
Published: (2025) -
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024) -
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023) -
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)