HGRN2: Gated Linear RNNs with State Expansion
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913471695683584 |
|---|---|
| author | Qin, Zhen Yang, Songlin Sun, Weixuan Shen, Xuyang Li, Dong Sun, Weigao Zhong, Yiran |
| author_facet | Qin, Zhen Yang, Songlin Sun, Weixuan Shen, Xuyang Li, Dong Sun, Weigao Zhong, Yiran |
| contents | Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However, the recurrent state size of HGRN remains relatively small, limiting its expressiveness. To address this issue, we introduce a simple outer product-based state expansion mechanism, which significantly enlarges the recurrent state size without introducing any additional parameters. This enhancement also provides a linear attention interpretation for HGRN2, enabling hardware-efficient training. Our extensive experiments verify the advantage of HGRN2 over HGRN consistently across different settings and competitive with other recurrent models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_07904 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | HGRN2: Gated Linear RNNs with State Expansion Qin, Zhen Yang, Songlin Sun, Weixuan Shen, Xuyang Li, Dong Sun, Weigao Zhong, Yiran Computation and Language Hierarchically gated linear RNN (HGRN, \citealt{HGRN}) has demonstrated competitive training speed and performance in language modeling while offering efficient inference. However, the recurrent state size of HGRN remains relatively small, limiting its expressiveness. To address this issue, we introduce a simple outer product-based state expansion mechanism, which significantly enlarges the recurrent state size without introducing any additional parameters. This enhancement also provides a linear attention interpretation for HGRN2, enabling hardware-efficient training. Our extensive experiments verify the advantage of HGRN2 over HGRN consistently across different settings and competitive with other recurrent models. |
| title | HGRN2: Gated Linear RNNs with State Expansion |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2404.07904 |