OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fernández-Hernández, Alberto, Mestre, Jose I., Pérez-Corral, Cristian, Dolz, Manuel F., Duato, Jose, Quintana-Ortí, Enrique S.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916000180469760
author Fernández-Hernández, Alberto
Mestre, Jose I.
Pérez-Corral, Cristian
Dolz, Manuel F.
Duato, Jose
Quintana-Ortí, Enrique S.
author_facet Fernández-Hernández, Alberto
Mestre, Jose I.
Pérez-Corral, Cristian
Dolz, Manuel F.
Duato, Jose
Quintana-Ortí, Enrique S.
contents Weight decay remains one of the most widely used regularization mechanisms for training convolutional neural networks, yet it is still commonly applied as a fixed coefficient shared by all layers throughout training. This uniform treatment ignores that different layers may follow different structural dynamics and therefore may require different regularization strengths. In this work, we propose OUIDecay, an adaptive layer-wise and time-dependent weight decay scheduler for CNNs driven by the Overfitting-Underfitting Indicator (OUI), an activation-based metric previously shown to provide early information about regularization quality. OUIDecay uses a lightweight batch-based formulation of OUI to monitor the structural behavior of each layer online and periodically rescales its weight decay relative to the other layers in the network. Unlike gradient-based adaptive decay methods, our approach relies on functional information extracted from activation patterns and does not require validation data. Experiments on EfficientNet-B0 with Stanford Cars, ResNet50 with Food101, DenseNet121 with CIFAR100, and MobileNetV2 with CIFAR10 show that OUIDecay achieves the best mean best-validation-loss in 7 out of 8 evaluated settings. These results indicate that activation-driven weight decay adaptation is a practical and effective alternative to fixed decay and gradient-based adaptive decay, while keeping the method lightweight and suitable for online use.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10161
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
Fernández-Hernández, Alberto
Mestre, Jose I.
Pérez-Corral, Cristian
Dolz, Manuel F.
Duato, Jose
Quintana-Ortí, Enrique S.
Machine Learning
Weight decay remains one of the most widely used regularization mechanisms for training convolutional neural networks, yet it is still commonly applied as a fixed coefficient shared by all layers throughout training. This uniform treatment ignores that different layers may follow different structural dynamics and therefore may require different regularization strengths. In this work, we propose OUIDecay, an adaptive layer-wise and time-dependent weight decay scheduler for CNNs driven by the Overfitting-Underfitting Indicator (OUI), an activation-based metric previously shown to provide early information about regularization quality. OUIDecay uses a lightweight batch-based formulation of OUI to monitor the structural behavior of each layer online and periodically rescales its weight decay relative to the other layers in the network. Unlike gradient-based adaptive decay methods, our approach relies on functional information extracted from activation patterns and does not require validation data. Experiments on EfficientNet-B0 with Stanford Cars, ResNet50 with Food101, DenseNet121 with CIFAR100, and MobileNetV2 with CIFAR10 show that OUIDecay achieves the best mean best-validation-loss in 7 out of 8 evaluated settings. These results indicate that activation-driven weight decay adaptation is a practical and effective alternative to fixed decay and gradient-based adaptive decay, while keeping the method lightweight and suitable for online use.
title OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
topic Machine Learning
url https://arxiv.org/abs/2605.10161