HGRN2: Gated Linear RNNs with State Expansion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Zhen, Yang, Songlin, Sun, Weixuan, Shen, Xuyang, Li, Dong, Sun, Weigao, Zhong, Yiran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913471695683584
author Qin, Zhen
Yang, Songlin
Sun, Weixuan
Shen, Xuyang
Li, Dong
Sun, Weigao
Zhong, Yiran
author_facet Qin, Zhen
Yang, Songlin
Sun, Weixuan
Shen, Xuyang
Li, Dong
Sun, Weigao
Zhong, Yiran
contents Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However, the recurrent state size of HGRN remains relatively small, limiting its expressiveness. To address this issue, we introduce a simple outer product-based state expansion mechanism, which significantly enlarges the recurrent state size without introducing any additional parameters. This enhancement also provides a linear attention interpretation for HGRN2, enabling hardware-efficient training. Our extensive experiments verify the advantage of HGRN2 over HGRN consistently across different settings and competitive with other recurrent models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07904
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HGRN2: Gated Linear RNNs with State Expansion
Qin, Zhen
Yang, Songlin
Sun, Weixuan
Shen, Xuyang
Li, Dong
Sun, Weigao
Zhong, Yiran
Computation and Language
Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However, the recurrent state size of HGRN remains relatively small, limiting its expressiveness. To address this issue, we introduce a simple outer product-based state expansion mechanism, which significantly enlarges the recurrent state size without introducing any additional parameters. This enhancement also provides a linear attention interpretation for HGRN2, enabling hardware-efficient training. Our extensive experiments verify the advantage of HGRN2 over HGRN consistently across different settings and competitive with other recurrent models.
title HGRN2: Gated Linear RNNs with State Expansion
topic Computation and Language
url https://arxiv.org/abs/2404.07904