Gradient Boosting within a Single Attention Layer
Fuente:
arXiv
Guardado en:
| Autor principal: | Sargolzaei, Saleh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
por: Zixian, Wang
Publicado: (2025)
por: Zixian, Wang
Publicado: (2025)
Gradient Boosting Reinforcement Learning
por: Fuhrer, Benjamin, et al.
Publicado: (2024)
por: Fuhrer, Benjamin, et al.
Publicado: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024)
por: Sheen, Heejune, et al.
Publicado: (2024)
AXIL: Exact Instance Attribution for Gradient Boosting
por: Geertsema, Paul, et al.
Publicado: (2023)
por: Geertsema, Paul, et al.
Publicado: (2023)
Supervised Score-Based Modeling by Gradient Boosting
por: Zhao, Changyuan, et al.
Publicado: (2024)
por: Zhao, Changyuan, et al.
Publicado: (2024)
Mixture of Layers with Hybrid Attention
por: Ternovtsii, Ivan, et al.
Publicado: (2026)
por: Ternovtsii, Ivan, et al.
Publicado: (2026)
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
por: Guan, Lei, et al.
Publicado: (2023)
por: Guan, Lei, et al.
Publicado: (2023)
UTBoost: Gradient Boosted Decision Trees for Uplift Modeling
por: Gao, Junjie, et al.
Publicado: (2023)
por: Gao, Junjie, et al.
Publicado: (2023)
FPBoost: Fully Parametric Gradient Boosting for Survival Analysis
por: Archetti, Alberto, et al.
Publicado: (2024)
por: Archetti, Alberto, et al.
Publicado: (2024)
SecureBoost+: Large Scale and High-Performance Vertical Federated Gradient Boosting Decision Tree
por: Fan, Tao, et al.
Publicado: (2021)
por: Fan, Tao, et al.
Publicado: (2021)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
por: Chen, Yihong, et al.
Publicado: (2026)
por: Chen, Yihong, et al.
Publicado: (2026)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
por: Chen, Hung-Hsuan
Publicado: (2024)
por: Chen, Hung-Hsuan
Publicado: (2024)
Universal Approximation Theorem for a Single-Layer Transformer
por: Gumaan, Esmail
Publicado: (2025)
por: Gumaan, Esmail
Publicado: (2025)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
por: Nguyen, Minh Duong, et al.
Publicado: (2024)
por: Nguyen, Minh Duong, et al.
Publicado: (2024)
Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
por: Tanna, Aditya, et al.
Publicado: (2026)
por: Tanna, Aditya, et al.
Publicado: (2026)
X-SAM: Boosting Sharpness-Aware Minimization with Dominant-Eigenvector Gradient Correction
por: Duan, Hongru, et al.
Publicado: (2026)
por: Duan, Hongru, et al.
Publicado: (2026)
PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting
por: Zhao, Bowen, et al.
Publicado: (2025)
por: Zhao, Bowen, et al.
Publicado: (2025)
Improving Out-of-Distribution Data Handling and Corruption Resistance via Modern Hopfield Networks
por: Sargolzaei, Saleh, et al.
Publicado: (2024)
por: Sargolzaei, Saleh, et al.
Publicado: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
por: Joshi, Sahil, et al.
Publicado: (2025)
por: Joshi, Sahil, et al.
Publicado: (2025)
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
por: Agarwal, Naman, et al.
Publicado: (2025)
por: Agarwal, Naman, et al.
Publicado: (2025)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
por: Datta, Shrestha, et al.
Publicado: (2026)
por: Datta, Shrestha, et al.
Publicado: (2026)
An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient Boosting and Fuzzy Rule-Based Models
por: Li, Jinbo, et al.
Publicado: (2025)
por: Li, Jinbo, et al.
Publicado: (2025)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
por: Zhang, Qixin, et al.
Publicado: (2024)
por: Zhang, Qixin, et al.
Publicado: (2024)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
por: Gavito, Andrea Treviño, et al.
Publicado: (2023)
por: Gavito, Andrea Treviño, et al.
Publicado: (2023)
Attention Smoothing Is All You Need For Unlearning
por: Zade, Saleh Zare, et al.
Publicado: (2026)
por: Zade, Saleh Zare, et al.
Publicado: (2026)
DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
por: Lin, Ci, et al.
Publicado: (2025)
por: Lin, Ci, et al.
Publicado: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
A generalized decision tree ensemble based on the NeuralNetworks architecture: Distributed Gradient Boosting Forest (DGBF)
por: Delgado-Panadero, Ángel, et al.
Publicado: (2024)
por: Delgado-Panadero, Ángel, et al.
Publicado: (2024)
MambaSL: Exploring Single-Layer Mamba for Time Series Classification
por: Jung, Yoo-Min, et al.
Publicado: (2026)
por: Jung, Yoo-Min, et al.
Publicado: (2026)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
por: Wu, Qitian, et al.
Publicado: (2024)
por: Wu, Qitian, et al.
Publicado: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
por: Kim, Hyunjun
Publicado: (2026)
por: Kim, Hyunjun
Publicado: (2026)
Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
por: Fartale, Harshwardhan, et al.
Publicado: (2025)
por: Fartale, Harshwardhan, et al.
Publicado: (2025)
AttriReBoost: A Gradient-Free Propagation Optimization Method for Cold Start Mitigation in Attribute Missing Graphs
por: Li, Mengran, et al.
Publicado: (2025)
por: Li, Mengran, et al.
Publicado: (2025)
SANGRIA: Stacked Autoencoder Neural Networks with Gradient Boosting for Indoor Localization
por: Gufran, Danish, et al.
Publicado: (2024)
por: Gufran, Danish, et al.
Publicado: (2024)
Instruction Following by Principled Boosting Attention of Large Language Models
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
por: Acero, Fernando, et al.
Publicado: (2024)
por: Acero, Fernando, et al.
Publicado: (2024)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
por: Riera, Carlos Boned, et al.
Publicado: (2025)
por: Riera, Carlos Boned, et al.
Publicado: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
por: Jeffares, Alan, et al.
Publicado: (2024)
por: Jeffares, Alan, et al.
Publicado: (2024)
GFPack++: Improving 2D Irregular Packing by Learning Gradient Field with Attention
por: Xue, Tianyang, et al.
Publicado: (2024)
por: Xue, Tianyang, et al.
Publicado: (2024)
Ejemplares similares
-
Attention Saturation and Gradient Suppression at Inflection Layers: Diagnosing and Mitigating Bottlenecks in Transformer Adaptation
por: Zixian, Wang
Publicado: (2025) -
Gradient Boosting Reinforcement Learning
por: Fuhrer, Benjamin, et al.
Publicado: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024) -
AXIL: Exact Instance Attribution for Gradient Boosting
por: Geertsema, Paul, et al.
Publicado: (2023) -
Supervised Score-Based Modeling by Gradient Boosting
por: Zhao, Changyuan, et al.
Publicado: (2024)