MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yulong, Liu, Xiang, Huang, Hongxiang, Lin, Xiaopeng, Liu, Zunchang, Chu, Xiaowen, Xie, Zeke, Cheng, Bojun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRF: Parallel Resonate and Fire Neuron for Long Sequence Learning in Spiking Neural Networks
by: Huang, Yulong, et al.
Published: (2024)
by: Huang, Yulong, et al.
Published: (2024)
CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks
by: Huang, Yulong, et al.
Published: (2024)
by: Huang, Yulong, et al.
Published: (2024)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
A Framework for Non-Linear Attention via Modern Hopfield Networks
by: Farooq, Ahmed
Published: (2025)
by: Farooq, Ahmed
Published: (2025)
Biologically Plausible Training of Deep Neural Networks Using a Top-down Credit Assignment Network
by: Chen, Jian-Hui, et al.
Published: (2022)
by: Chen, Jian-Hui, et al.
Published: (2022)
Dynamically Weighted Momentum with Adaptive Step Sizes for Efficient Deep Network Training
by: Wang, Zhifeng, et al.
Published: (2025)
by: Wang, Zhifeng, et al.
Published: (2025)
Spiking Heterogeneous Graph Attention Networks
by: Cao, Buqing, et al.
Published: (2025)
by: Cao, Buqing, et al.
Published: (2025)
Natural Evolutionary Search meets Probabilistic Numerics
by: Osselin, Pierre, et al.
Published: (2025)
by: Osselin, Pierre, et al.
Published: (2025)
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
by: Dong, Pingcheng, et al.
Published: (2024)
by: Dong, Pingcheng, et al.
Published: (2024)
Equivariant Neural Networks for General Linear Symmetries on Lie Algebras
by: Kim, Chankyo, et al.
Published: (2025)
by: Kim, Chankyo, et al.
Published: (2025)
Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training
by: Luo, Xiaoling, et al.
Published: (2025)
by: Luo, Xiaoling, et al.
Published: (2025)
Exploiting Chaotic Dynamics as Deep Neural Networks
by: Liu, Shuhong, et al.
Published: (2024)
by: Liu, Shuhong, et al.
Published: (2024)
Bullet Trains: Parallelizing Training of Temporally Precise Spiking Neural Networks
by: Morrill, Todd, et al.
Published: (2026)
by: Morrill, Todd, et al.
Published: (2026)
Multiple Population Alternate Evolution Neural Architecture Search
by: Zou, Juan, et al.
Published: (2024)
by: Zou, Juan, et al.
Published: (2024)
Expanded Gating Ranges Improve Activation Functions
by: Huang, Allen Hao
Published: (2024)
by: Huang, Allen Hao
Published: (2024)
SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
by: Huang, Yulong, et al.
Published: (2025)
by: Huang, Yulong, et al.
Published: (2025)
Towards Initialization-Agnostic Clustering with Iterative Adaptive Resonance Theory
by: Qu, Xiaozheng, et al.
Published: (2025)
by: Qu, Xiaozheng, et al.
Published: (2025)
Forward Direct Feedback Alignment for Online Gradient Estimates of Spiking Neural Networks
by: Bacho, Florian, et al.
Published: (2024)
by: Bacho, Florian, et al.
Published: (2024)
HEATACO: Heatmap-Guided Ant Colony Decoding for Large-Scale Travelling Salesman Problems
by: Lin, Bo-Cheng, et al.
Published: (2026)
by: Lin, Bo-Cheng, et al.
Published: (2026)
Deriving Activation Functions Using Integration
by: Huang, Allen Hao, et al.
Published: (2024)
by: Huang, Allen Hao, et al.
Published: (2024)
Robust MAE-Driven NAS: From Mask Reconstruction to Architecture Innovation
by: Hu, Yiming, et al.
Published: (2023)
by: Hu, Yiming, et al.
Published: (2023)
Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity
by: Pierro, Alessandro, et al.
Published: (2025)
by: Pierro, Alessandro, et al.
Published: (2025)
Spatiotemporal Forecasting of Traffic Flow using Wavelet-based Temporal Attention
by: Jakhmola, Yash, et al.
Published: (2024)
by: Jakhmola, Yash, et al.
Published: (2024)
Data-Informed Model Complexity Metric for Optimizing Symbolic Regression Models
by: Haut, Nathan, et al.
Published: (2025)
by: Haut, Nathan, et al.
Published: (2025)
Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data Clustering
by: Zhang, Yiqun, et al.
Published: (2024)
by: Zhang, Yiqun, et al.
Published: (2024)
SpikeGraphormer: A High-Performance Graph Transformer with Spiking Graph Attention
by: Sun, Yundong, et al.
Published: (2024)
by: Sun, Yundong, et al.
Published: (2024)
Hierarchical Kernel Transformer: Multi-Scale Attention with an Information-Theoretic Approximation Analysis
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Research on short-term load forecasting model based on VMD and IPSO-ELM
by: Xie, Qiang
Published: (2024)
by: Xie, Qiang
Published: (2024)
Linearly Constrained Weights: Reducing Activation Shift for Faster Training of Neural Networks
by: Kutsuna, Takuro
Published: (2024)
by: Kutsuna, Takuro
Published: (2024)
ReLiCADA -- Reservoir Computing using Linear Cellular Automata Design Algorithm
by: Kantic, Jonas, et al.
Published: (2023)
by: Kantic, Jonas, et al.
Published: (2023)
Stable and Robust Deep Learning By Hyperbolic Tangent Exponential Linear Unit (TeLU)
by: Fernandez, Alfredo, et al.
Published: (2024)
by: Fernandez, Alfredo, et al.
Published: (2024)
Noisy Spiking Actor Network for Exploration
by: Chen, Ding, et al.
Published: (2024)
by: Chen, Ding, et al.
Published: (2024)
GridPE: Unifying Positional Encoding in Transformers with a Grid Cell-Inspired Framework
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)
by: Orvieto, Antonio, et al.
Published: (2023)
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
by: Zhang, Huizhe, et al.
Published: (2025)
by: Zhang, Huizhe, et al.
Published: (2025)
NiSNN-A: Non-iterative Spiking Neural Networks with Attention with Application to Motor Imagery EEG Classification
by: Zhang, Chuhan, et al.
Published: (2023)
by: Zhang, Chuhan, et al.
Published: (2023)
RsGCN: Subgraph-Based Rescaling Enhances Generalization of GCNs for Solving Traveling Salesman Problems
by: Huang, Junquan, et al.
Published: (2025)
by: Huang, Junquan, et al.
Published: (2025)
Neural networks for neurocomputing circuits: a computational study of tolerance to noise and activation function non-uniformity when machine learning materials properties
by: Thant, Ye min, et al.
Published: (2025)
by: Thant, Ye min, et al.
Published: (2025)
READY: Reward Discovery for Meta-Black-Box Optimization
by: Huang, Zechuan, et al.
Published: (2026)
by: Huang, Zechuan, et al.
Published: (2026)
Similar Items
-
PRF: Parallel Resonate and Fire Neuron for Long Sequence Learning in Spiking Neural Networks
by: Huang, Yulong, et al.
Published: (2024) -
CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks
by: Huang, Yulong, et al.
Published: (2024) -
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024) -
A Framework for Non-Linear Attention via Modern Hopfield Networks
by: Farooq, Ahmed
Published: (2025) -
Biologically Plausible Training of Deep Neural Networks Using a Top-down Credit Assignment Network
by: Chen, Jian-Hui, et al.
Published: (2022)