Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization
Fuente:
arXiv
Saved in:
| Main Author: | Kavun, Sergii |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loss shaping enhances exact gradient learning with Eventprop in spiking neural networks
by: Nowotny, Thomas, et al.
Published: (2022)
by: Nowotny, Thomas, et al.
Published: (2022)
Model Input-Output Configuration Search with Embedded Feature Selection for Sensor Time-series and Image Classification
by: Hoang, Anh T., et al.
Published: (2023)
by: Hoang, Anh T., et al.
Published: (2023)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
On the Equivalence of Regression and Classification
by: Jayadeva, et al.
Published: (2025)
by: Jayadeva, et al.
Published: (2025)
Spiking Sequence Machines and Transformers
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
Graded Transformers
by: Shaska Sr, Tony
Published: (2025)
by: Shaska Sr, Tony
Published: (2025)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
by: Sharma, Paras
Published: (2026)
by: Sharma, Paras
Published: (2026)
Convexity-Driven Projection for Point Cloud Dimensionality Reduction
by: Sanyal, Suman
Published: (2025)
by: Sanyal, Suman
Published: (2025)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
by: Štrupl, Miroslav, et al.
Published: (2025)
by: Štrupl, Miroslav, et al.
Published: (2025)
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
by: Alagöz, Celal, et al.
Published: (2026)
by: Alagöz, Celal, et al.
Published: (2026)
Retrieval-Augmented Memory for Online Learning
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
An in-depth look at approximation via deep and narrow neural networks
by: Dommel, Joris, et al.
Published: (2025)
by: Dommel, Joris, et al.
Published: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
by: BC, Samiksha
Published: (2025)
by: BC, Samiksha
Published: (2025)
The Coordinate System Problem in Persistent Structural Memory for Neural Architectures
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026)
by: Agrawal, Vidhi, et al.
Published: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
Energy-Efficient Neuromorphic Computing for Edge AI: A Framework with Adaptive Spiking Neural Networks and Hardware-Aware Optimization
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
The Geometry of Cortical Computation: Manifold Disentanglement and Predictive Dynamics in VCNet
by: Hill, Brennen A., et al.
Published: (2025)
by: Hill, Brennen A., et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
Benchmarking changepoint detection algorithms on cardiac time series
by: Cakmak, Ayse, et al.
Published: (2024)
by: Cakmak, Ayse, et al.
Published: (2024)
Deep Spectral Meshes: Multi-Frequency Facial Mesh Processing with Graph Neural Networks
by: Kosk, Robert, et al.
Published: (2024)
by: Kosk, Robert, et al.
Published: (2024)
Perforated Neural Networks for Keyword Spotting
by: Gopal, Vishy, et al.
Published: (2026)
by: Gopal, Vishy, et al.
Published: (2026)
mlx-snn: Spiking Neural Networks on Apple Silicon via MLX
by: Qin, Jiahao
Published: (2026)
by: Qin, Jiahao
Published: (2026)
Diverse capability and scaling of diffusion and auto-regressive models when learning abstract rules
by: Wang, Binxu, et al.
Published: (2024)
by: Wang, Binxu, et al.
Published: (2024)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Distal Interference: Exploring the Limits of Model-Based Continual Learning
by: van Deventer, Heinrich, et al.
Published: (2024)
by: van Deventer, Heinrich, et al.
Published: (2024)
An arithmetic method algorithm optimizing k-nearest neighbors compared to regression algorithms and evaluated on real world data sources
by: Anagnostopoulos, Theodoros, et al.
Published: (2026)
by: Anagnostopoulos, Theodoros, et al.
Published: (2026)
Energy-Efficient Information Representation in MNIST Classification Using Biologically Inspired Learning
by: Stricker, Patrick, et al.
Published: (2026)
by: Stricker, Patrick, et al.
Published: (2026)
NOVAK: Unified adaptive optimizer for deep neural networks
by: Kavun, Sergii
Published: (2026)
by: Kavun, Sergii
Published: (2026)
Randomized Spline Trees for Functional Data Classification: Theory and Application to Environmental Time Series
by: Riccio, Donato, et al.
Published: (2024)
by: Riccio, Donato, et al.
Published: (2024)
Similar Items
-
Loss shaping enhances exact gradient learning with Eventprop in spiking neural networks
by: Nowotny, Thomas, et al.
Published: (2022) -
Model Input-Output Configuration Search with Embedded Feature Selection for Sensor Time-series and Image Classification
by: Hoang, Anh T., et al.
Published: (2023) -
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025) -
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025) -
On the Equivalence of Regression and Classification
by: Jayadeva, et al.
Published: (2025)