Customizing the Inductive Biases of Softmax Attention using Structured Matrices
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Yilun, Amsel, Noah, Lotfi, Sanae, Qiu, Shikai, Potapczynski, Andres, Wilson, Andrew Gordon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
by: Potapczynski, Andres, et al.
Published: (2024)
by: Potapczynski, Andres, et al.
Published: (2024)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
by: Lotfi, Sanae, et al.
Published: (2024)
by: Lotfi, Sanae, et al.
Published: (2024)
Non-Vacuous Generalization Bounds for Large Language Models
by: Lotfi, Sanae, et al.
Published: (2023)
by: Lotfi, Sanae, et al.
Published: (2023)
Training Flexible Models of Genetic Variant Effects from Functional Annotations using Accelerated Linear Algebra
by: Amin, Alan N., et al.
Published: (2025)
by: Amin, Alan N., et al.
Published: (2025)
On the Benefits of Rank in Attention Layers
by: Amsel, Noah, et al.
Published: (2024)
by: Amsel, Noah, et al.
Published: (2024)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025)
by: Marek, Martin, et al.
Published: (2025)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
by: Goldblum, Micah, et al.
Published: (2023)
by: Goldblum, Micah, et al.
Published: (2023)
Large Language Models Are Zero-Shot Time Series Forecasters
by: Gruver, Nate, et al.
Published: (2023)
by: Gruver, Nate, et al.
Published: (2023)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Dion2: A Simple Method to Shrink Matrix in Muon
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
by: Yehudai, Gilad, et al.
Published: (2025)
by: Yehudai, Gilad, et al.
Published: (2025)
Transformers Are Born Biased: Structural Inductive Biases at Random Initialization and Their Practical Consequences
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)
by: Marek, Martin, et al.
Published: (2026)
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
by: Finzi, Marc, et al.
Published: (2026)
by: Finzi, Marc, et al.
Published: (2026)
The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases
by: Romero, David W.
Published: (2024)
by: Romero, David W.
Published: (2024)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Instilling Inductive Biases with Subnetworks
by: Zhang, Enyan, et al.
Published: (2023)
by: Zhang, Enyan, et al.
Published: (2023)
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Characterising the Inductive Biases of Neural Networks on Boolean Data
by: Mingard, Chris, et al.
Published: (2025)
by: Mingard, Chris, et al.
Published: (2025)
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling
by: Mittal, Daksh, et al.
Published: (2025)
by: Mittal, Daksh, et al.
Published: (2025)
Incorporating Inductive Biases to Energy-based Generative Models
by: Li, Yukun, et al.
Published: (2025)
by: Li, Yukun, et al.
Published: (2025)
Theoretical Analysis of Inductive Biases in Deep Convolutional Networks
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Clustering Inductive Biases with Unrolled Networks
by: Huml, Jonathan, et al.
Published: (2023)
by: Huml, Jonathan, et al.
Published: (2023)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
by: Lubana, Ekdeep Singh, et al.
Published: (2025)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
by: Movahedi, Sajad, et al.
Published: (2024)
by: Movahedi, Sajad, et al.
Published: (2024)
Deep Learning is Not So Mysterious or Different
by: Wilson, Andrew Gordon
Published: (2025)
by: Wilson, Andrew Gordon
Published: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Out-of-Distribution Detection Methods Answer the Wrong Questions
by: Li, Yucen Lily, et al.
Published: (2025)
by: Li, Yucen Lily, et al.
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Diffusion Model's Generalization Can Be Characterized by Inductive Biases toward a Data-Dependent Ridge Manifold
by: He, Ye, et al.
Published: (2026)
by: He, Ye, et al.
Published: (2026)
Similar Items
-
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024) -
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
by: Potapczynski, Andres, et al.
Published: (2024) -
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
by: Lotfi, Sanae, et al.
Published: (2024) -
Non-Vacuous Generalization Bounds for Large Language Models
by: Lotfi, Sanae, et al.
Published: (2023) -
Training Flexible Models of Genetic Variant Effects from Functional Annotations using Accelerated Linear Algebra
by: Amin, Alan N., et al.
Published: (2025)