Contrast-Enhanced Gating in GRUs for Robust Low-Data Sequence Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Subramanian, Barathi, Jeyaraj, Rathinaraja, Paul, Anand
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914511466790912
author Subramanian, Barathi
Jeyaraj, Rathinaraja
Paul, Anand
author_facet Subramanian, Barathi
Jeyaraj, Rathinaraja
Paul, Anand
contents Activation functions govern how recurrent networks regulate and transmit information across temporal dependencies. Despite advances in sequence modelling, gated recurrent units (GRUs) still depend on the standard sigmoid and tanh nonlinearities, which can produce weak gate separation and unstable learning, particularly when training data are limited. We introduce squared sigmoid-tanh (SST), a parameter-free activation that squares the gate nonlinearity to increase contrast between near-zero- and high-activations, thereby promoting sharper information filtering during GRU updates. We incorporate SST into GRU gating and evaluate it across low-data settings spanning sign language recognition, human activity recognition, and time-series forecasting and classification. Across tasks, SST-GRU consistently surpasses standard sigmoid/tanh GRU, with the largest improvements observed in the smallest-data domains, while adding negligible computational cost. We further examine gate activation statistics and training dynamics, showing that SST improves training stability, which aligns with its performance gains in data-scarce settings. SST is a parameter-free modification that complements more complex architectural advances by improving gating selectivity in low-data sequence learning.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09034
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrast-Enhanced Gating in GRUs for Robust Low-Data Sequence Learning
Subramanian, Barathi
Jeyaraj, Rathinaraja
Paul, Anand
Machine Learning
Artificial Intelligence
Activation functions govern how recurrent networks regulate and transmit information across temporal dependencies. Despite advances in sequence modelling, gated recurrent units (GRUs) still depend on the standard sigmoid and tanh nonlinearities, which can produce weak gate separation and unstable learning, particularly when training data are limited. We introduce squared sigmoid-tanh (SST), a parameter-free activation that squares the gate nonlinearity to increase contrast between near-zero- and high-activations, thereby promoting sharper information filtering during GRU updates. We incorporate SST into GRU gating and evaluate it across low-data settings spanning sign language recognition, human activity recognition, and time-series forecasting and classification. Across tasks, SST-GRU consistently surpasses standard sigmoid/tanh GRU, with the largest improvements observed in the smallest-data domains, while adding negligible computational cost. We further examine gate activation statistics and training dynamics, showing that SST improves training stability, which aligns with its performance gains in data-scarce settings. SST is a parameter-free modification that complements more complex architectural advances by improving gating selectivity in low-data sequence learning.
title Contrast-Enhanced Gating in GRUs for Robust Low-Data Sequence Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.09034