Softlog-Softmax Layers and Divergences Contribute to a Computationally Dependable Ensemble Learning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Atto, Abdourrahmane Mahamane |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
von: Salis, Matteo, et al.
Veröffentlicht: (2024)
von: Salis, Matteo, et al.
Veröffentlicht: (2024)
Pure and Physics-Guided Deep Learning Solutions for Spatio-Temporal Groundwater Level Prediction at Arbitrary Locations
von: Salis, Matteo, et al.
Veröffentlicht: (2026)
von: Salis, Matteo, et al.
Veröffentlicht: (2026)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
Smooth SCAD: A Raised Cosine SCAD Type Thresholding Rule for Wavelet Denoising
von: Kulkarni, Radhika, et al.
Veröffentlicht: (2026)
von: Kulkarni, Radhika, et al.
Veröffentlicht: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
Online Continual Learning via Logit Adjusted Softmax
von: Huang, Zhehao, et al.
Veröffentlicht: (2023)
von: Huang, Zhehao, et al.
Veröffentlicht: (2023)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
von: Sheen, Heejune, et al.
Veröffentlicht: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
Softmax as a Lagrangian-Legendrian Seam
von: Lee-Jenkins, Christopher R.
Veröffentlicht: (2025)
von: Lee-Jenkins, Christopher R.
Veröffentlicht: (2025)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
von: Pu, Yuanhao, et al.
Veröffentlicht: (2026)
von: Pu, Yuanhao, et al.
Veröffentlicht: (2026)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
von: Rezazadeh, Navid, et al.
Veröffentlicht: (2026)
von: Rezazadeh, Navid, et al.
Veröffentlicht: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
von: Kühn, Marcel, et al.
Veröffentlicht: (2026)
von: Kühn, Marcel, et al.
Veröffentlicht: (2026)
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
von: Cheng, Hailing, et al.
Veröffentlicht: (2026)
von: Cheng, Hailing, et al.
Veröffentlicht: (2026)
Divergent Ensemble Networks: Enhancing Uncertainty Estimation with Shared Representations and Independent Branching
von: Kharbanda, Arnav, et al.
Veröffentlicht: (2024)
von: Kharbanda, Arnav, et al.
Veröffentlicht: (2024)
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
von: Shi, Zhongjie, et al.
Veröffentlicht: (2026)
von: Shi, Zhongjie, et al.
Veröffentlicht: (2026)
Universal Approximation with Softmax Attention
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Gumbel-MPNN: Graph Rewiring with Gumbel-Softmax
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2025)
von: Hoffmann, Marcel, et al.
Veröffentlicht: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2025)
von: Nguyen, Huy, et al.
Veröffentlicht: (2025)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
von: Pu, Yuanhao, et al.
Veröffentlicht: (2025)
von: Pu, Yuanhao, et al.
Veröffentlicht: (2025)
CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
von: Park, Hanjun, et al.
Veröffentlicht: (2026)
von: Park, Hanjun, et al.
Veröffentlicht: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Towards Understanding Layer Contributions in Tabular In-Context Learning Models
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
von: Yan, Fanqi, et al.
Veröffentlicht: (2025)
Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
von: Zuhri, Zayd M. K., et al.
Veröffentlicht: (2025)
von: Zuhri, Zayd M. K., et al.
Veröffentlicht: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025)
von: Wang, Run, et al.
Veröffentlicht: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
von: Attar, Navid Akhavan, et al.
Veröffentlicht: (2026)
von: Attar, Navid Akhavan, et al.
Veröffentlicht: (2026)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
von: Ebrahimpour-Boroojeny, Ali, et al.
Veröffentlicht: (2024)
von: Ebrahimpour-Boroojeny, Ali, et al.
Veröffentlicht: (2024)
Scalable-Softmax Is Superior for Attention
von: Nakanishi, Ken M.
Veröffentlicht: (2025)
von: Nakanishi, Ken M.
Veröffentlicht: (2025)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2025)
von: Mongaras, Gabriel, et al.
Veröffentlicht: (2025)
Beyond Softmax: A Natural Parameterization for Categorical Random Variables
von: Manenti, Alessandro, et al.
Veröffentlicht: (2025)
von: Manenti, Alessandro, et al.
Veröffentlicht: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
von: Kuang, Yilun, et al.
Veröffentlicht: (2025)
von: Kuang, Yilun, et al.
Veröffentlicht: (2025)
Binary Hypothesis Testing for Softmax Models and Leverage Score Models
von: Gu, Yuzhou, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhou, et al.
Veröffentlicht: (2024)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
von: Lattimore, Tor
Veröffentlicht: (2026)
von: Lattimore, Tor
Veröffentlicht: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
von: Collins, Liam, et al.
Veröffentlicht: (2024)
von: Collins, Liam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Time Distributed Deep Learning Models for Purely Exogenous Forecasting: Application to Water Table Depth Prediction using Weather Image Time Series
von: Salis, Matteo, et al.
Veröffentlicht: (2024) -
Pure and Physics-Guided Deep Learning Solutions for Spatio-Temporal Groundwater Level Prediction at Arbitrary Locations
von: Salis, Matteo, et al.
Veröffentlicht: (2026) -
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
von: Lv, Qi, et al.
Veröffentlicht: (2025) -
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024) -
Smooth SCAD: A Raised Cosine SCAD Type Thresholding Rule for Wavelet Denoising
von: Kulkarni, Radhika, et al.
Veröffentlicht: (2026)