FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Pingwei, Hu, Yuxuan, Tan, Jianchao, Wang, Xue, Zhang, Jiaqi, Lu, Yifan, Sun, Yerui, Xie, Yuchen, Cai, Xunliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917456192208896
author Sun, Pingwei
Hu, Yuxuan
Tan, Jianchao
Wang, Xue
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
author_facet Sun, Pingwei
Hu, Yuxuan
Tan, Jianchao
Wang, Xue
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
contents Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent advances such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA) have demonstrated that the delta rule, an online gradient descent update, enables superior associative recall compared to simple additive updates. While KDA refined the coarse head-wise decay gate into channel-wise decay, the learning rate $β_t$ in the delta update remains a scalar, limiting the model's capacity for dimension-specific adaptation. We introduce FG$^2$-GDN, which replaces the scalar $β_t$ with a channel-wise vector analogous to the transition from SGD to per-coordinate adaptive optimizers such as AdaGrad and Adam. We further propose FG$^2$-GDN+, which decouples the scaling for keys and values, enabling independent control of erasure strength and write strength. Experiments on synthetic and real-world benchmarks show that FG$^2$-GDN and its variant improve associative recall and long-context understanding over GDN and KDA, with comparable computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
Sun, Pingwei
Hu, Yuxuan
Tan, Jianchao
Wang, Xue
Zhang, Jiaqi
Lu, Yifan
Sun, Yerui
Xie, Yuchen
Cai, Xunliang
Machine Learning
Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent advances such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA) have demonstrated that the delta rule, an online gradient descent update, enables superior associative recall compared to simple additive updates. While KDA refined the coarse head-wise decay gate into channel-wise decay, the learning rate $β_t$ in the delta update remains a scalar, limiting the model's capacity for dimension-specific adaptation. We introduce FG$^2$-GDN, which replaces the scalar $β_t$ with a channel-wise vector analogous to the transition from SGD to per-coordinate adaptive optimizers such as AdaGrad and Adam. We further propose FG$^2$-GDN+, which decouples the scaling for keys and values, enabling independent control of erasure strength and write strength. Experiments on synthetic and real-world benchmarks show that FG$^2$-GDN and its variant improve associative recall and long-context understanding over GDN and KDA, with comparable computational efficiency.
title FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
topic Machine Learning
url https://arxiv.org/abs/2604.19021