State Rank Dynamics in Linear Attention LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Ao, Zhang, Hongtao, Zhou, Heng, Ma, Yixuan, Qin, Yiran, Su, Tongrui, Liu, Yan, Ma, Zhanyu, Xu, Jun, Gao, Jiuchong, Hao, Jinghua, He, Renqing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910008860475392
author Sun, Ao
Zhang, Hongtao
Zhou, Heng
Ma, Yixuan
Qin, Yiran
Su, Tongrui
Liu, Yan
Ma, Zhanyu
Xu, Jun
Gao, Jiuchong
Hao, Jinghua
He, Renqing
author_facet Sun, Ao
Zhang, Hongtao
Zhou, Heng
Ma, Yixuan
Qin, Yiran
Su, Tongrui
Liu, Yan
Ma, Zhanyu
Xu, Jun
Gao, Jiuchong
Hao, Jinghua
He, Renqing
contents Linear Attention Large Language Models (LLMs) offer a compelling recurrent formulation that compresses context into a fixed-size state matrix, enabling constant-time inference. However, the internal dynamics of this compressed state remain largely opaque. In this work, we present a comprehensive study on the runtime state dynamics of state-of-the-art Linear Attention models. We uncover a fundamental phenomenon termed State Rank Stratification, characterized by a distinct spectral bifurcation among linear attention heads: while one group maintains an effective rank oscillating near zero, the other exhibits rapid growth that converges to an upper bound. Extensive experiments across diverse inference contexts reveal that these dynamics remain strikingly consistent, indicating that the identity of a head,whether low-rank or high-rank,is an intrinsic structural property acquired during pre-training, rather than a transient state dependent on the input data. Furthermore, our diagnostic probes reveal a surprising functional divergence: low-rank heads are indispensable for model reasoning, whereas high-rank heads exhibit significant redundancy. Leveraging this insight, we propose Joint Rank-Norm Pruning, a zero-shot strategy that achieves a 38.9\% reduction in KV-cache overhead while largely maintaining model accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle State Rank Dynamics in Linear Attention LLMs
Sun, Ao
Zhang, Hongtao
Zhou, Heng
Ma, Yixuan
Qin, Yiran
Su, Tongrui
Liu, Yan
Ma, Zhanyu
Xu, Jun
Gao, Jiuchong
Hao, Jinghua
He, Renqing
Machine Learning
Artificial Intelligence
Linear Attention Large Language Models (LLMs) offer a compelling recurrent formulation that compresses context into a fixed-size state matrix, enabling constant-time inference. However, the internal dynamics of this compressed state remain largely opaque. In this work, we present a comprehensive study on the runtime state dynamics of state-of-the-art Linear Attention models. We uncover a fundamental phenomenon termed State Rank Stratification, characterized by a distinct spectral bifurcation among linear attention heads: while one group maintains an effective rank oscillating near zero, the other exhibits rapid growth that converges to an upper bound. Extensive experiments across diverse inference contexts reveal that these dynamics remain strikingly consistent, indicating that the identity of a head,whether low-rank or high-rank,is an intrinsic structural property acquired during pre-training, rather than a transient state dependent on the input data. Furthermore, our diagnostic probes reveal a surprising functional divergence: low-rank heads are indispensable for model reasoning, whereas high-rank heads exhibit significant redundancy. Leveraging this insight, we propose Joint Rank-Norm Pruning, a zero-shot strategy that achieves a 38.9\% reduction in KV-cache overhead while largely maintaining model accuracy.
title State Rank Dynamics in Linear Attention LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.02195