Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Saratchandran, Hemanth, Zheng, Jianqiao, Ji, Yiping, Zhang, Wenbo, Lucey, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)
by: Ji, Yiping, et al.
Published: (2024)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Enhancing Transformers Through Conditioned Embedded Tokens
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
by: Gordon, Cameron, et al.
Published: (2024)
by: Gordon, Cameron, et al.
Published: (2024)
SineProject: Machine Unlearning for Stable Vision Language Alignment
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Spectral Conditioning of Attention Improves Transformer Performance
by: Saratchandran, Hemanth, et al.
Published: (2026)
by: Saratchandran, Hemanth, et al.
Published: (2026)
Weight Conditioning for Smooth Optimization of Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
by: Zheng, Jianqiao, et al.
Published: (2026)
by: Zheng, Jianqiao, et al.
Published: (2026)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
Rethinking the Role of Spatial Mixing
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Invertible Neural Warp for NeRF
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Convolutional Initialization for Data-Efficient Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Gradient Descent as a Shrinkage Operator for Spectral Bias
by: Lucey, Simon
Published: (2025)
by: Lucey, Simon
Published: (2025)
The Quantization Benefits of Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2026)
by: Ji, Yiping, et al.
Published: (2026)
Cutting the Skip: Training Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026)
by: Saratchandran, Hemanth
Published: (2026)
Trading Positional Complexity vs. Deepness in Coordinate Networks
by: Zheng, Jianqiao, et al.
Published: (2022)
by: Zheng, Jianqiao, et al.
Published: (2022)
$ε$-Softmax: Approximating One-Hot Vectors for Mitigating Label Noise
by: Wang, Jialiang, et al.
Published: (2025)
by: Wang, Jialiang, et al.
Published: (2025)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023)
by: Hoffmann, David T., et al.
Published: (2023)
SineLoRA$Δ$: Sine-Activated Delta Compression
by: Gordon, Cameron, et al.
Published: (2025)
by: Gordon, Cameron, et al.
Published: (2025)
What Does Softmax Probability Tell Us about Classifiers Ranking Across Diverse Test Conditions?
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
3D-LFM: Lifting Foundation Model
by: Dabhi, Mosam, et al.
Published: (2023)
by: Dabhi, Mosam, et al.
Published: (2023)
Rethink Arbitrary Style Transfer with Transformer and Contrastive Learning
by: Zhang, Zhanjie, et al.
Published: (2024)
by: Zhang, Zhanjie, et al.
Published: (2024)
Explorations of the Softmax Space: Knowing When the Neural Network Doesn't Know
by: Sikar, Daniel, et al.
Published: (2025)
by: Sikar, Daniel, et al.
Published: (2025)
Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
by: Wang, Jeffrey, et al.
Published: (2026)
by: Wang, Jeffrey, et al.
Published: (2026)
Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
by: Oh, Yoojin, et al.
Published: (2025)
by: Oh, Yoojin, et al.
Published: (2025)
Analyzing the Neural Tangent Kernel of Periodically Activated Coordinate Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Architectural Strategies for the optimization of Physics-Informed Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Analytical Softmax Temperature Setting from Feature Dimensions for Model- and Domain-Robust Classification
by: Hasegawa, Tatsuhito, et al.
Published: (2025)
by: Hasegawa, Tatsuhito, et al.
Published: (2025)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026)
by: Song, Alan Z., et al.
Published: (2026)
Similar Items
-
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025) -
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025) -
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025) -
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025) -
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)