HGRN2: Gated Linear RNNs with State Expansion
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Zhen, Yang, Songlin, Sun, Weixuan, Shen, Xuyang, Li, Dong, Sun, Weigao, Zhong, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Scaling Laws for Linear Complexity Language Models
by: Shen, Xuyang, et al.
Published: (2024)
by: Shen, Xuyang, et al.
Published: (2024)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Unlocking the Secrets of Linear Complexity Sequence Model from A Unified Perspective
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Elucidating the Design Space of Decay in Linear Attention
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Comba: Improving Bilinear RNNs with Closed-loop Control
by: Hu, Jiaxi, et al.
Published: (2025)
by: Hu, Jiaxi, et al.
Published: (2025)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
by: Lan, Disen, et al.
Published: (2025)
by: Lan, Disen, et al.
Published: (2025)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Why Are Linear RNNs More Parallelizable?
by: Merrill, William, et al.
Published: (2026)
by: Merrill, William, et al.
Published: (2026)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Learning State-Tracking from Code Using Linear RNNs
by: Siems, Julien, et al.
Published: (2026)
by: Siems, Julien, et al.
Published: (2026)
Transformers are Multi-State RNNs
by: Oren, Matanel, et al.
Published: (2024)
by: Oren, Matanel, et al.
Published: (2024)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
by: Siems, Julien, et al.
Published: (2025)
by: Siems, Julien, et al.
Published: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
MoM: Linear Sequence Modeling with Mixture-of-Memories
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
Gated Delta Networks: Improving Mamba2 with Delta Rule
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
StateX: Enhancing RNN Recall via Post-training State Expansion
by: Shen, Xingyu, et al.
Published: (2025)
by: Shen, Xingyu, et al.
Published: (2025)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
by: Shen, Sherrie, et al.
Published: (2025)
by: Shen, Sherrie, et al.
Published: (2025)
On Efficiently Representing Regular Languages as RNNs
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Uncovering the Computational Roles of Nonlinearity in Sequence Modeling Using Almost-Linear RNNs
by: Brenner, Manuel, et al.
Published: (2025)
by: Brenner, Manuel, et al.
Published: (2025)
Native Hybrid Attention for Efficient Sequence Modeling
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
by: Wang, Weixuan, et al.
Published: (2024)
by: Wang, Weixuan, et al.
Published: (2024)
TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning
by: Sun, Bowen, et al.
Published: (2026)
by: Sun, Bowen, et al.
Published: (2026)
RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
by: Shen, Yiting, et al.
Published: (2026)
by: Shen, Yiting, et al.
Published: (2026)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Similar Items
-
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
by: Qin, Zhen, et al.
Published: (2024) -
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
by: Qin, Zhen, et al.
Published: (2024) -
Scaling Laws for Linear Complexity Language Models
by: Shen, Xuyang, et al.
Published: (2024) -
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024) -
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)