Spectral Conditioning of Attention Improves Transformer Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Saratchandran, Hemanth, Lucey, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026)
by: Saratchandran, Hemanth
Published: (2026)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
by: Zheng, Jianqiao, et al.
Published: (2026)
by: Zheng, Jianqiao, et al.
Published: (2026)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Enhancing Transformers Through Conditioned Embedded Tokens
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Analyzing the Neural Tangent Kernel of Periodically Activated Coordinate Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Architectural Strategies for the optimization of Physics-Informed Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
The Quantization Benefits of Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2026)
by: Ji, Yiping, et al.
Published: (2026)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
A Sampling Theory Perspective on Activations for Implicit Neural Representations
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
SineLoRA$Δ$: Sine-Activated Delta Compression
by: Gordon, Cameron, et al.
Published: (2025)
by: Gordon, Cameron, et al.
Published: (2025)
D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
by: Gordon, Cameron, et al.
Published: (2024)
by: Gordon, Cameron, et al.
Published: (2024)
Cutting the Skip: Training Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)
by: Ji, Yiping, et al.
Published: (2024)
Weight Conditioning for Smooth Optimization of Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
SineProject: Machine Unlearning for Stable Vision Language Alignment
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Gradient Descent as a Shrinkage Operator for Spectral Bias
by: Lucey, Simon
Published: (2025)
by: Lucey, Simon
Published: (2025)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Data Denoising and Derivative Estimation for Data-Driven Modeling of Nonlinear Dynamical Systems
by: Yao, Jiaqi, et al.
Published: (2025)
by: Yao, Jiaqi, et al.
Published: (2025)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
Invertible Neural Warp for NeRF
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
by: Horton, Michael, et al.
Published: (2025)
by: Horton, Michael, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
A Spectral Condition for Feature Learning
by: Yang, Greg, et al.
Published: (2023)
by: Yang, Greg, et al.
Published: (2023)
Rethinking the Role of Spatial Mixing
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Neuromorphic Mimicry Attacks Exploiting Brain-Inspired Computing for Covert Cyber Intrusions
by: Ravipati, Hemanth
Published: (2025)
by: Ravipati, Hemanth
Published: (2025)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
Selective Attention Improves Transformer
by: Leviathan, Yaniv, et al.
Published: (2024)
by: Leviathan, Yaniv, et al.
Published: (2024)
A Ruelle dynamical zeta function for equivariant flows
by: Hochs, Peter, et al.
Published: (2023)
by: Hochs, Peter, et al.
Published: (2023)
An equivariant Guillemin trace formula
by: Hochs, Peter, et al.
Published: (2025)
by: Hochs, Peter, et al.
Published: (2025)
Fed-Meta-Align: A Similarity-Aware Aggregation and Personalization Pipeline for Federated TinyML on Heterogeneous Data
by: Macharla, Hemanth, et al.
Published: (2025)
by: Macharla, Hemanth, et al.
Published: (2025)
Spectral Energy Centroid: a Metric for Improving Performance and Analyzing Spectral Bias in Implicit Neural Representations
by: Dądela, Tomasz, et al.
Published: (2026)
by: Dądela, Tomasz, et al.
Published: (2026)
Advancing Household Robotics: Deep Interactive Reinforcement Learning for Efficient Training and Enhanced Performance
by: Soni, Arpita, et al.
Published: (2024)
by: Soni, Arpita, et al.
Published: (2024)
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
Similar Items
-
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026) -
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025) -
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
by: Zheng, Jianqiao, et al.
Published: (2026) -
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025) -
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)