Explicit Dropout: Deterministic Regularization for Transformer Architectures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Agrawal, Vidhi, Oleksiienko, Illia, Iosifidis, Alexandros
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908986881605632
author Agrawal, Vidhi
Oleksiienko, Illia
Iosifidis, Alexandros
author_facet Agrawal, Vidhi
Oleksiienko, Illia
Iosifidis, Alexandros
contents Dropout is a widely used regularization technique in deep learning, but its effects are typically realized through stochastic masking rather than explicit optimization objectives. We propose a deterministic formulation that expresses dropout as an additive regularizer directly incorporated into the training loss. The framework derives explicit regularization terms for Transformer architectures, covering attention query, key, value, and feed-forward components with independently controllable strengths. This formulation removes reliance on stochastic perturbations while providing clearer and fine-grained control over regularization strength. Experiments across image classification, temporal action detection, and audio classification show that explicit dropout matches or outperforms conventional implicit methods, with consistent gains when applied to attention and feed-forward network layers. Ablation studies demonstrate stable performance and controllable regularization through regularization coefficients and dropout rates. Overall, explicit dropout offers a practical and interpretable alternative to stochastic regularization while maintaining architectural flexibility across diverse tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20505
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explicit Dropout: Deterministic Regularization for Transformer Architectures
Agrawal, Vidhi
Oleksiienko, Illia
Iosifidis, Alexandros
Machine Learning
68T07
I.2.6; I.5.1
Dropout is a widely used regularization technique in deep learning, but its effects are typically realized through stochastic masking rather than explicit optimization objectives. We propose a deterministic formulation that expresses dropout as an additive regularizer directly incorporated into the training loss. The framework derives explicit regularization terms for Transformer architectures, covering attention query, key, value, and feed-forward components with independently controllable strengths. This formulation removes reliance on stochastic perturbations while providing clearer and fine-grained control over regularization strength. Experiments across image classification, temporal action detection, and audio classification show that explicit dropout matches or outperforms conventional implicit methods, with consistent gains when applied to attention and feed-forward network layers. Ablation studies demonstrate stable performance and controllable regularization through regularization coefficients and dropout rates. Overall, explicit dropout offers a practical and interpretable alternative to stochastic regularization while maintaining architectural flexibility across diverse tasks.
title Explicit Dropout: Deterministic Regularization for Transformer Architectures
topic Machine Learning
68T07
I.2.6; I.5.1
url https://arxiv.org/abs/2604.20505