Transformer tricks: Precomputing the first layer
Fuente:
arXiv
Saved in:
| Main Author: | Graef, Nils |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KV-weights are all you need for skipless transformers
by: Graef, Nils
Published: (2024)
by: Graef, Nils
Published: (2024)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
by: Graef, Nils, et al.
Published: (2025)
by: Graef, Nils, et al.
Published: (2025)
FlashNorm: Fast Normalization for Transformers
by: Graef, Nils, et al.
Published: (2024)
by: Graef, Nils, et al.
Published: (2024)
Fitting networks with a cancellation trick
by: Jin, Jiashun, et al.
Published: (2025)
by: Jin, Jiashun, et al.
Published: (2025)
Generalized vec trick for fast learning of pairwise kernel models
by: Viljanen, Markus, et al.
Published: (2020)
by: Viljanen, Markus, et al.
Published: (2020)
Reciprocal Latent Fields for Precomputed Sound Propagation
by: Seuté, Hugo, et al.
Published: (2026)
by: Seuté, Hugo, et al.
Published: (2026)
Signs of the Past, Patterns of the Present: On the Automatic Classification of Old Babylonian Cuneiform Signs
by: Verwimp, Eli, et al.
Published: (2025)
by: Verwimp, Eli, et al.
Published: (2025)
First line of defense: A robust first layer mitigates adversarial attacks
by: Suresh, Janani, et al.
Published: (2024)
by: Suresh, Janani, et al.
Published: (2024)
PDE-Transformer: Efficient and Versatile Transformers for Physics Simulations
by: Holzschuh, Benjamin, et al.
Published: (2025)
by: Holzschuh, Benjamin, et al.
Published: (2025)
Approximating invariant functions with the sorting trick is theoretically justified
by: Chaimanowong, Wee, et al.
Published: (2024)
by: Chaimanowong, Wee, et al.
Published: (2024)
Teaming LLMs to Detect and Mitigate Hallucinations
by: Till, Demian, et al.
Published: (2025)
by: Till, Demian, et al.
Published: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
Insights Into the Inner Workings of Transformer Models for Protein Function Prediction
by: Wenzel, Markus, et al.
Published: (2023)
by: Wenzel, Markus, et al.
Published: (2023)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Deep-ICE: the first globally optimal algorithm for minimizing 0-1 loss in two-layer ReLU and maxout networks
by: He, Xi, et al.
Published: (2025)
by: He, Xi, et al.
Published: (2025)
Theoretical limitations of multi-layer Transformer
by: Chen, Lijie, et al.
Published: (2024)
by: Chen, Lijie, et al.
Published: (2024)
Toward generalizable learning of all (linear) first-order methods via memory augmented Transformers
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
On the Unreasonable Effectiveness of Last-layer Retraining
by: Hill, John C., et al.
Published: (2025)
by: Hill, John C., et al.
Published: (2025)
The Unreasonable Effectiveness of Solving Inverse Problems with Neural Networks
by: Holl, Philipp, et al.
Published: (2024)
by: Holl, Philipp, et al.
Published: (2024)
Flow Matching for Posterior Inference with Simulator Feedback
by: Holzschuh, Benjamin, et al.
Published: (2024)
by: Holzschuh, Benjamin, et al.
Published: (2024)
Neural Emulator Superiority: When Machine Learning for PDEs Surpasses its Training Data
by: Koehler, Felix, et al.
Published: (2025)
by: Koehler, Felix, et al.
Published: (2025)
Supervised Training Rapidly Degrades Early Visual Cortex Alignment Across Biologically Plausible Learning Rules
by: Leutenegger, Nils
Published: (2026)
by: Leutenegger, Nils
Published: (2026)
Myosotis: structured computation for attention like layer
by: Egorov, Evgenii, et al.
Published: (2025)
by: Egorov, Evgenii, et al.
Published: (2025)
Stochastic Gradient Descent for Two-layer Neural Networks
by: Cao, Dinghao, et al.
Published: (2024)
by: Cao, Dinghao, et al.
Published: (2024)
Latent class analysis for multi-layer categorical data
by: Qing, Huan
Published: (2024)
by: Qing, Huan
Published: (2024)
Attention layers provably solve single-location regression
by: Marion, Pierre, et al.
Published: (2024)
by: Marion, Pierre, et al.
Published: (2024)
Tabular data generation with tensor contraction layers and transformers
by: Silva, Aníbal, et al.
Published: (2024)
by: Silva, Aníbal, et al.
Published: (2024)
Task agnostic continual learning with Pairwise layer architecture
by: Keskinen, Santtu
Published: (2024)
by: Keskinen, Santtu
Published: (2024)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Lower bounds for one-layer transformers that compute parity
by: Hsu, Daniel
Published: (2026)
by: Hsu, Daniel
Published: (2026)
Multi-layer Stack Ensembles for Time Series Forecasting
by: Bosch, Nathanael, et al.
Published: (2025)
by: Bosch, Nathanael, et al.
Published: (2025)
Simplifying Random Forests' Probabilistic Forecasts
by: Koster, Nils, et al.
Published: (2024)
by: Koster, Nils, et al.
Published: (2024)
Efficient Symbolic Computations for Identifying Causal Effects
by: Hollering, Benjamin, et al.
Published: (2026)
by: Hollering, Benjamin, et al.
Published: (2026)
PRDP: Progressively Refined Differentiable Physics
by: Bhatia, Kanishk, et al.
Published: (2025)
by: Bhatia, Kanishk, et al.
Published: (2025)
Conveyance: A Versatile Framework for Learning in Structured Class Spaces
by: Taha, Yasser, et al.
Published: (2026)
by: Taha, Yasser, et al.
Published: (2026)
Collaborative Threshold Watermarking
by: Bakr, Tameem, et al.
Published: (2026)
by: Bakr, Tameem, et al.
Published: (2026)
Injectivity of ReLU-layers: Tools from Frame Theory
by: Haider, Daniel, et al.
Published: (2024)
by: Haider, Daniel, et al.
Published: (2024)
One-layer transformers fail to solve the induction heads task
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
HOT: Higher-Order Dynamic Graph Representation Learning with Efficient Transformers
by: Besta, Maciej, et al.
Published: (2023)
by: Besta, Maciej, et al.
Published: (2023)
School Library Management.
by: Graef, Bob, et al.
Published: (1988)
by: Graef, Bob, et al.
Published: (1988)
Similar Items
-
KV-weights are all you need for skipless transformers
by: Graef, Nils
Published: (2024) -
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
by: Graef, Nils, et al.
Published: (2025) -
FlashNorm: Fast Normalization for Transformers
by: Graef, Nils, et al.
Published: (2024) -
Fitting networks with a cancellation trick
by: Jin, Jiashun, et al.
Published: (2025) -
Generalized vec trick for fast learning of pairwise kernel models
by: Viljanen, Markus, et al.
Published: (2020)