Gradient Boosting within a Single Attention Layer
Fuente:
arXiv
Saved in:
| Main Author: | Sargolzaei, Saleh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
AXIL: Exact Instance Attribution for Gradient Boosting
by: Geertsema, Paul, et al.
Published: (2023)
by: Geertsema, Paul, et al.
Published: (2023)
Supervised Score-Based Modeling by Gradient Boosting
by: Zhao, Changyuan, et al.
Published: (2024)
by: Zhao, Changyuan, et al.
Published: (2024)
Mixture of Layers with Hybrid Attention
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
by: Guan, Lei, et al.
Published: (2023)
by: Guan, Lei, et al.
Published: (2023)
UTBoost: Gradient Boosted Decision Trees for Uplift Modeling
by: Gao, Junjie, et al.
Published: (2023)
by: Gao, Junjie, et al.
Published: (2023)
FPBoost: Fully Parametric Gradient Boosting for Survival Analysis
by: Archetti, Alberto, et al.
Published: (2024)
by: Archetti, Alberto, et al.
Published: (2024)
SecureBoost+: Large Scale and High-Performance Vertical Federated Gradient Boosting Decision Tree
by: Fan, Tao, et al.
Published: (2021)
by: Fan, Tao, et al.
Published: (2021)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
by: Chen, Hung-Hsuan
Published: (2024)
by: Chen, Hung-Hsuan
Published: (2024)
Universal Approximation Theorem for a Single-Layer Transformer
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
by: Nguyen, Minh Duong, et al.
Published: (2024)
by: Nguyen, Minh Duong, et al.
Published: (2024)
Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
X-SAM: Boosting Sharpness-Aware Minimization with Dominant-Eigenvector Gradient Correction
by: Duan, Hongru, et al.
Published: (2026)
by: Duan, Hongru, et al.
Published: (2026)
PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
by: Zhao, Bowen, et al.
Published: (2025)
by: Zhao, Bowen, et al.
Published: (2025)
Improving Out-of-Distribution Data Handling and Corruption Resistance via Modern Hopfield Networks
by: Sargolzaei, Saleh, et al.
Published: (2024)
by: Sargolzaei, Saleh, et al.
Published: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)
by: Datta, Shrestha, et al.
Published: (2026)
An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient Boosting and Fuzzy Rule-Based Models
by: Li, Jinbo, et al.
Published: (2025)
by: Li, Jinbo, et al.
Published: (2025)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
by: Zhang, Qixin, et al.
Published: (2024)
by: Zhang, Qixin, et al.
Published: (2024)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
by: Gavito, Andrea Treviño, et al.
Published: (2023)
by: Gavito, Andrea Treviño, et al.
Published: (2023)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
by: Lin, Ci, et al.
Published: (2025)
by: Lin, Ci, et al.
Published: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
A generalized decision tree ensemble based on the NeuralNetworks architecture: Distributed Gradient Boosting Forest (DGBF)
by: Delgado-Panadero, Ángel, et al.
Published: (2024)
by: Delgado-Panadero, Ángel, et al.
Published: (2024)
MambaSL: Exploring Single-Layer Mamba for Time Series Classification
by: Jung, Yoo-Min, et al.
Published: (2026)
by: Jung, Yoo-Min, et al.
Published: (2026)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
by: Wu, Qitian, et al.
Published: (2024)
by: Wu, Qitian, et al.
Published: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
by: Fartale, Harshwardhan, et al.
Published: (2025)
by: Fartale, Harshwardhan, et al.
Published: (2025)
AttriReBoost: A Gradient-Free Propagation Optimization Method for Cold Start Mitigation in Attribute Missing Graphs
by: Li, Mengran, et al.
Published: (2025)
by: Li, Mengran, et al.
Published: (2025)
SANGRIA: Stacked Autoencoder Neural Networks with Gradient Boosting for Indoor Localization
by: Gufran, Danish, et al.
Published: (2024)
by: Gufran, Danish, et al.
Published: (2024)
Instruction Following by Principled Boosting Attention of Large Language Models
by: Guardieiro, Vitoria, et al.
Published: (2025)
by: Guardieiro, Vitoria, et al.
Published: (2025)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
by: Acero, Fernando, et al.
Published: (2024)
by: Acero, Fernando, et al.
Published: (2024)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
by: Riera, Carlos Boned, et al.
Published: (2025)
by: Riera, Carlos Boned, et al.
Published: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
by: Jeffares, Alan, et al.
Published: (2024)
by: Jeffares, Alan, et al.
Published: (2024)
GFPack++: Improving 2D Irregular Packing by Learning Gradient Field with Attention
by: Xue, Tianyang, et al.
Published: (2024)
by: Xue, Tianyang, et al.
Published: (2024)
Similar Items
-
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
by: Zixian, Wang
Published: (2025) -
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024) -
AXIL: Exact Instance Attribution for Gradient Boosting
by: Geertsema, Paul, et al.
Published: (2023) -
Supervised Score-Based Modeling by Gradient Boosting
by: Zhao, Changyuan, et al.
Published: (2024)