Explicit Dropout: Deterministic Regularization for Transformer Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Vidhi, Oleksiienko, Illia, Iosifidis, Alexandros |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
by: Medeiros, Daniel Nobrega
Published: (2026)
by: Medeiros, Daniel Nobrega
Published: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
by: Dai, Yanning, et al.
Published: (2026)
by: Dai, Yanning, et al.
Published: (2026)
Predicting Traffic Accident Severity with Deep Neural Networks
by: Bibb, Meghan, et al.
Published: (2025)
by: Bibb, Meghan, et al.
Published: (2025)
An in-depth look at approximation via deep and narrow neural networks
by: Dommel, Joris, et al.
Published: (2025)
by: Dommel, Joris, et al.
Published: (2025)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
Complex-Valued Phase-Coherent Transformer
by: Hioki, Leona
Published: (2026)
by: Hioki, Leona
Published: (2026)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
by: Sharma, Paras
Published: (2026)
by: Sharma, Paras
Published: (2026)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Revisiting GAN with Bayes-Optimal Discrimination
by: Naeini, Mohammadreza Tavasoli, et al.
Published: (2025)
by: Naeini, Mohammadreza Tavasoli, et al.
Published: (2025)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
JacNet: Learning Functions with Structured Jacobians
by: Lorraine, Jonathan, et al.
Published: (2024)
by: Lorraine, Jonathan, et al.
Published: (2024)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
by: Lau, Matthew, et al.
Published: (2023)
by: Lau, Matthew, et al.
Published: (2023)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
by: Di Santi, Eduardo, et al.
Published: (2025)
by: Di Santi, Eduardo, et al.
Published: (2025)
Step-E: A Differentiable Data Cleaning Framework for Robust Learning with Noisy Labels
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
GraphNNK -- Graph Classification and Interpretability
by: Bolevic, Zeljko, et al.
Published: (2026)
by: Bolevic, Zeljko, et al.
Published: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
RCUKF: Data-Driven Modeling Meets Bayesian Estimation
by: Anurag, Kumar, et al.
Published: (2025)
by: Anurag, Kumar, et al.
Published: (2025)
PH-VAE: A Polynomial Hierarchical Variational Autoencoder Towards Disentangled Representation Learning
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Estimating the Event-Related Potential from Few EEG Trials
by: Nørskov, Anders Vestergaard, et al.
Published: (2025)
by: Nørskov, Anders Vestergaard, et al.
Published: (2025)
Optimized Gradient Clipping for Noisy Label Learning
by: Ye, Xichen, et al.
Published: (2024)
by: Ye, Xichen, et al.
Published: (2024)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Mathematical Foundations of Neural Tangents and Infinite-Width Networks
by: Mysore, Rachana, et al.
Published: (2025)
by: Mysore, Rachana, et al.
Published: (2025)
Similar Items
-
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025) -
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026) -
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025) -
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025) -
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)