Gradient Boosting within a Single Attention Layer
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Sargolzaei, Saleh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
von: Zixian, Wang
Veröffentlicht: (2025)
von: Zixian, Wang
Veröffentlicht: (2025)
Gradient Boosting Reinforcement Learning
von: Fuhrer, Benjamin, et al.
Veröffentlicht: (2024)
von: Fuhrer, Benjamin, et al.
Veröffentlicht: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
AXIL: Exact Instance Attribution for Gradient Boosting
von: Geertsema, Paul, et al.
Veröffentlicht: (2023)
von: Geertsema, Paul, et al.
Veröffentlicht: (2023)
Supervised Score-Based Modeling by Gradient Boosting
von: Zhao, Changyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Changyuan, et al.
Veröffentlicht: (2024)
Mixture of Layers with Hybrid Attention
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2026)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
von: Guan, Lei, et al.
Veröffentlicht: (2023)
von: Guan, Lei, et al.
Veröffentlicht: (2023)
UTBoost: Gradient Boosted Decision Trees for Uplift Modeling
von: Gao, Junjie, et al.
Veröffentlicht: (2023)
von: Gao, Junjie, et al.
Veröffentlicht: (2023)
FPBoost: Fully Parametric Gradient Boosting for Survival Analysis
von: Archetti, Alberto, et al.
Veröffentlicht: (2024)
von: Archetti, Alberto, et al.
Veröffentlicht: (2024)
SecureBoost+: Large Scale and High-Performance Vertical Federated Gradient Boosting Decision Tree
von: Fan, Tao, et al.
Veröffentlicht: (2021)
von: Fan, Tao, et al.
Veröffentlicht: (2021)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
von: Chen, Yihong, et al.
Veröffentlicht: (2026)
von: Chen, Yihong, et al.
Veröffentlicht: (2026)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
von: Chen, Hung-Hsuan
Veröffentlicht: (2024)
von: Chen, Hung-Hsuan
Veröffentlicht: (2024)
Universal Approximation Theorem for a Single-Layer Transformer
von: Gumaan, Esmail
Veröffentlicht: (2025)
von: Gumaan, Esmail
Veröffentlicht: (2025)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
von: Nguyen, Minh Duong, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh Duong, et al.
Veröffentlicht: (2024)
Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
von: Tanna, Aditya, et al.
Veröffentlicht: (2026)
von: Tanna, Aditya, et al.
Veröffentlicht: (2026)
X-SAM: Boosting Sharpness-Aware Minimization with Dominant-Eigenvector Gradient Correction
von: Duan, Hongru, et al.
Veröffentlicht: (2026)
von: Duan, Hongru, et al.
Veröffentlicht: (2026)
PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
von: Zhao, Bowen, et al.
Veröffentlicht: (2025)
von: Zhao, Bowen, et al.
Veröffentlicht: (2025)
Improving Out-of-Distribution Data Handling and Corruption Resistance via Modern Hopfield Networks
von: Sargolzaei, Saleh, et al.
Veröffentlicht: (2024)
von: Sargolzaei, Saleh, et al.
Veröffentlicht: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
von: Saikumar, Dhananjay, et al.
Veröffentlicht: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
von: Datta, Shrestha, et al.
Veröffentlicht: (2026)
von: Datta, Shrestha, et al.
Veröffentlicht: (2026)
An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient Boosting and Fuzzy Rule-Based Models
von: Li, Jinbo, et al.
Veröffentlicht: (2025)
von: Li, Jinbo, et al.
Veröffentlicht: (2025)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
von: Gavito, Andrea Treviño, et al.
Veröffentlicht: (2023)
von: Gavito, Andrea Treviño, et al.
Veröffentlicht: (2023)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
von: Lin, Ci, et al.
Veröffentlicht: (2025)
von: Lin, Ci, et al.
Veröffentlicht: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
A generalized decision tree ensemble based on the NeuralNetworks architecture: Distributed Gradient Boosting Forest (DGBF)
von: Delgado-Panadero, Ángel, et al.
Veröffentlicht: (2024)
von: Delgado-Panadero, Ángel, et al.
Veröffentlicht: (2024)
MambaSL: Exploring Single-Layer Mamba for Time Series Classification
von: Jung, Yoo-Min, et al.
Veröffentlicht: (2026)
von: Jung, Yoo-Min, et al.
Veröffentlicht: (2026)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
von: Wu, Qitian, et al.
Veröffentlicht: (2024)
von: Wu, Qitian, et al.
Veröffentlicht: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
von: Kim, Hyunjun
Veröffentlicht: (2026)
von: Kim, Hyunjun
Veröffentlicht: (2026)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
von: Fartale, Harshwardhan, et al.
Veröffentlicht: (2025)
von: Fartale, Harshwardhan, et al.
Veröffentlicht: (2025)
AttriReBoost: A Gradient-Free Propagation Optimization Method for Cold Start Mitigation in Attribute Missing Graphs
von: Li, Mengran, et al.
Veröffentlicht: (2025)
von: Li, Mengran, et al.
Veröffentlicht: (2025)
SANGRIA: Stacked Autoencoder Neural Networks with Gradient Boosting for Indoor Localization
von: Gufran, Danish, et al.
Veröffentlicht: (2024)
von: Gufran, Danish, et al.
Veröffentlicht: (2024)
Instruction Following by Principled Boosting Attention of Large Language Models
von: Guardieiro, Vitoria, et al.
Veröffentlicht: (2025)
von: Guardieiro, Vitoria, et al.
Veröffentlicht: (2025)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
von: Acero, Fernando, et al.
Veröffentlicht: (2024)
von: Acero, Fernando, et al.
Veröffentlicht: (2024)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
von: Jeffares, Alan, et al.
Veröffentlicht: (2024)
von: Jeffares, Alan, et al.
Veröffentlicht: (2024)
GFPack++: Improving 2D Irregular Packing by Learning Gradient Field with Attention
von: Xue, Tianyang, et al.
Veröffentlicht: (2024)
von: Xue, Tianyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
von: Zixian, Wang
Veröffentlicht: (2025) -
Gradient Boosting Reinforcement Learning
von: Fuhrer, Benjamin, et al.
Veröffentlicht: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024) -
AXIL: Exact Instance Attribution for Gradient Boosting
von: Geertsema, Paul, et al.
Veröffentlicht: (2023) -
Supervised Score-Based Modeling by Gradient Boosting
von: Zhao, Changyuan, et al.
Veröffentlicht: (2024)