Tangma: A Tanh-Guided Activation Function with Learnable Parameters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Golwala, Shreel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912484036706304
author Golwala, Shreel
author_facet Golwala, Shreel
contents Activation functions are key to effective backpropagation and expressiveness in deep neural networks. This work introduces Tangma, a new activation function that combines the smooth shape of the hyperbolic tangent with two learnable parameters: $α$, which shifts the curve's inflection point to adjust neuron activation, and $γ$, which adds linearity to preserve weak gradients and improve training stability. Tangma was evaluated on MNIST and CIFAR-10 using custom networks composed of convolutional and linear layers, and compared against ReLU, Swish, and GELU. On MNIST, Tangma achieved the highest validation accuracy of 99.09% and the lowest validation loss, demonstrating faster and more stable convergence than the baselines. On CIFAR-10, Tangma reached a top validation accuracy of 78.15%, outperforming all other activation functions while maintaining a competitive training loss. Tangma also showed improved training efficiency, with lower average epoch runtimes compared to Swish and GELU. These results suggest that Tangma performs well on standard vision tasks and enables reliable, efficient training. Its learnable design gives more control over activation behavior, which may benefit larger models in tasks such as image recognition or language modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10560
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tangma: A Tanh-Guided Activation Function with Learnable Parameters
Golwala, Shreel
Neural and Evolutionary Computing
Computer Vision and Pattern Recognition
Machine Learning
Activation functions are key to effective backpropagation and expressiveness in deep neural networks. This work introduces Tangma, a new activation function that combines the smooth shape of the hyperbolic tangent with two learnable parameters: $α$, which shifts the curve's inflection point to adjust neuron activation, and $γ$, which adds linearity to preserve weak gradients and improve training stability. Tangma was evaluated on MNIST and CIFAR-10 using custom networks composed of convolutional and linear layers, and compared against ReLU, Swish, and GELU. On MNIST, Tangma achieved the highest validation accuracy of 99.09% and the lowest validation loss, demonstrating faster and more stable convergence than the baselines. On CIFAR-10, Tangma reached a top validation accuracy of 78.15%, outperforming all other activation functions while maintaining a competitive training loss. Tangma also showed improved training efficiency, with lower average epoch runtimes compared to Swish and GELU. These results suggest that Tangma performs well on standard vision tasks and enables reliable, efficient training. Its learnable design gives more control over activation behavior, which may benefit larger models in tasks such as image recognition or language modeling.
title Tangma: A Tanh-Guided Activation Function with Learnable Parameters
topic Neural and Evolutionary Computing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.10560