ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
Fuente:
arXiv
Saved in:
| Main Author: | BC, Samiksha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025)
by: Bolshim, Maxim, et al.
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
by: Kassinos, Stavros C.
Published: (2025)
by: Kassinos, Stavros C.
Published: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
by: Jiang, Shuai, et al.
Published: (2026)
by: Jiang, Shuai, et al.
Published: (2026)
Deep Legendre Transform
by: Minabutdinov, Aleksey, et al.
Published: (2025)
by: Minabutdinov, Aleksey, et al.
Published: (2025)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization
by: Kavun, Sergii
Published: (2025)
by: Kavun, Sergii
Published: (2025)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026)
by: Agrawal, Vidhi, et al.
Published: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
Retrieval-Augmented Memory for Online Learning
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
by: Zhang, Minxin, et al.
Published: (2026)
by: Zhang, Minxin, et al.
Published: (2026)
Graded Transformers
by: Shaska Sr, Tony
Published: (2025)
by: Shaska Sr, Tony
Published: (2025)
ConvXformer: Differentially Private Hybrid ConvNeXt-Transformer for Inertial Navigation
by: Tariq, Omer, et al.
Published: (2025)
by: Tariq, Omer, et al.
Published: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
by: Lotfi, Ali, et al.
Published: (2026)
by: Lotfi, Ali, et al.
Published: (2026)
TED++: Submanifold-Aware Backdoor Detection via Layerwise Tubular-Neighbourhood Screening
by: Le, Nam, et al.
Published: (2025)
by: Le, Nam, et al.
Published: (2025)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
Three-dimensional inversion of gravity data using implicit neural representations and scientific machine learning
by: Mishra, Pankaj K, et al.
Published: (2025)
by: Mishra, Pankaj K, et al.
Published: (2025)
JacNet: Learning Functions with Structured Jacobians
by: Lorraine, Jonathan, et al.
Published: (2024)
by: Lorraine, Jonathan, et al.
Published: (2024)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
NeurOptimisation: The Spiking Way to Evolve
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
i-DEQ: A stable inertial deep equilibrium model for image restoration
by: Clerc, Antonin, et al.
Published: (2026)
by: Clerc, Antonin, et al.
Published: (2026)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
Seed-Induced Uniqueness in Transformer Models: Subspace Alignment Governs Subliminal Transfer
by: Okatan, Ayşe Selin, et al.
Published: (2025)
by: Okatan, Ayşe Selin, et al.
Published: (2025)
Mathematical Foundations of Neural Tangents and Infinite-Width Networks
by: Mysore, Rachana, et al.
Published: (2025)
by: Mysore, Rachana, et al.
Published: (2025)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Predicting Traffic Accident Severity with Deep Neural Networks
by: Bibb, Meghan, et al.
Published: (2025)
by: Bibb, Meghan, et al.
Published: (2025)
Similar Items
-
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025) -
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026) -
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
by: Bolshim, Maxim, et al.
Published: (2026) -
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025) -
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
by: Du, Wenzhang
Published: (2025)