Softlog-Softmax Layers and Divergences Contribute to a Computationally Dependable Ensemble Learning
Fuente:
arXiv
Saved in:
| Main Author: | Atto, Abdourrahmane Mahamane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
by: Salis, Matteo, et al.
Published: (2024)
by: Salis, Matteo, et al.
Published: (2024)
Pure and Physics-Guided Deep Learning Solutions for Spatio-Temporal Groundwater Level Prediction at Arbitrary Locations
by: Salis, Matteo, et al.
Published: (2026)
by: Salis, Matteo, et al.
Published: (2026)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Smooth SCAD: A Raised Cosine SCAD Type Thresholding Rule for Wavelet Denoising
by: Kulkarni, Radhika, et al.
Published: (2026)
by: Kulkarni, Radhika, et al.
Published: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
by: Choi, Dawon, et al.
Published: (2026)
by: Choi, Dawon, et al.
Published: (2026)
Online Continual Learning via Logit Adjusted Softmax
by: Huang, Zhehao, et al.
Published: (2023)
by: Huang, Zhehao, et al.
Published: (2023)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Softmax as a Lagrangian-Legendrian Seam
by: Lee-Jenkins, Christopher R.
Published: (2025)
by: Lee-Jenkins, Christopher R.
Published: (2025)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
by: Wang, Wenxun, et al.
Published: (2025)
by: Wang, Wenxun, et al.
Published: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026)
by: Kühn, Marcel, et al.
Published: (2026)
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
by: Cheng, Hailing, et al.
Published: (2026)
by: Cheng, Hailing, et al.
Published: (2026)
Divergent Ensemble Networks: Enhancing Uncertainty Estimation with Shared Representations and Independent Branching
by: Kharbanda, Arnav, et al.
Published: (2024)
by: Kharbanda, Arnav, et al.
Published: (2024)
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
by: Shi, Zhongjie, et al.
Published: (2026)
by: Shi, Zhongjie, et al.
Published: (2026)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
by: Hoffmann, Marcel, et al.
Published: (2025)
by: Hoffmann, Marcel, et al.
Published: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
by: Park, Hanjun, et al.
Published: (2026)
by: Park, Hanjun, et al.
Published: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Towards Understanding Layer Contributions in Tabular In-Context Learning Models
by: Balef, Amir Rezaei, et al.
Published: (2025)
by: Balef, Amir Rezaei, et al.
Published: (2025)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025)
by: Wang, Run, et al.
Published: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2024)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
by: Mongaras, Gabriel, et al.
Published: (2025)
by: Mongaras, Gabriel, et al.
Published: (2025)
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
by: Manenti, Alessandro, et al.
Published: (2025)
by: Manenti, Alessandro, et al.
Published: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Binary Hypothesis Testing for Softmax Models and Leverage Score Models
by: Gu, Yuzhou, et al.
Published: (2024)
by: Gu, Yuzhou, et al.
Published: (2024)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Similar Items
-
Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
by: Salis, Matteo, et al.
Published: (2024) -
Pure and Physics-Guided Deep Learning Solutions for Spatio-Temporal Groundwater Level Prediction at Arbitrary Locations
by: Salis, Matteo, et al.
Published: (2026) -
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025) -
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024) -
Smooth SCAD: A Raised Cosine SCAD Type Thresholding Rule for Wavelet Denoising
by: Kulkarni, Radhika, et al.
Published: (2026)