Efficient Low Rank Attention for Long-Context Inference in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tenghui, Zhou, Guoxu, Zhao, Xuyang, Qiu, Yuning, Zhao, Qibin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918260081950720
author Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Qiu, Yuning
Zhao, Qibin
author_facet Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Qiu, Yuning
Zhao, Qibin
contents As the length of input text increases, the key-value (KV) cache in LLMs imposes prohibitive GPU memory costs and limits long-context inference on resource constrained devices. Existing approaches, such as KV quantization and pruning, reduce memory usage but suffer from numerical precision loss or suboptimal retention of key-value pairs. In this work, Low Rank Query and Key attention (LRQK) is introduced, a two-stage framework that jointly decomposes full-precision query and key matrices into compact rank-\(r\) factors during the prefill stage, and then employs these low-dimensional projections to compute proxy attention scores in \(\mathcal{O}(lr)\) time at each decode step. By selecting only the top-\(k\) tokens and a small fixed set of recent tokens, LRQK employs a mixed GPU-CPU cache with a hit-and-miss mechanism where only missing full-precision KV pairs are transferred, thereby preserving exact attention outputs while reducing CPU-GPU data movement. Extensive experiments on the RULER and LongBench benchmarks with LLaMA-3-8B and Qwen2.5-7B demonstrate that LRQK matches or surpasses leading sparse-attention methods in long context settings, while delivering significant memory savings with minimal accuracy loss. Our code is available at https://github.com/tenghuilee/LRQK.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Low Rank Attention for Long-Context Inference in Large Language Models
Li, Tenghui
Zhou, Guoxu
Zhao, Xuyang
Qiu, Yuning
Zhao, Qibin
Machine Learning
Artificial Intelligence
As the length of input text increases, the key-value (KV) cache in LLMs imposes prohibitive GPU memory costs and limits long-context inference on resource constrained devices. Existing approaches, such as KV quantization and pruning, reduce memory usage but suffer from numerical precision loss or suboptimal retention of key-value pairs. In this work, Low Rank Query and Key attention (LRQK) is introduced, a two-stage framework that jointly decomposes full-precision query and key matrices into compact rank-\(r\) factors during the prefill stage, and then employs these low-dimensional projections to compute proxy attention scores in \(\mathcal{O}(lr)\) time at each decode step. By selecting only the top-\(k\) tokens and a small fixed set of recent tokens, LRQK employs a mixed GPU-CPU cache with a hit-and-miss mechanism where only missing full-precision KV pairs are transferred, thereby preserving exact attention outputs while reducing CPU-GPU data movement. Extensive experiments on the RULER and LongBench benchmarks with LLaMA-3-8B and Qwen2.5-7B demonstrate that LRQK matches or surpasses leading sparse-attention methods in long context settings, while delivering significant memory savings with minimal accuracy loss. Our code is available at https://github.com/tenghuilee/LRQK.
title Efficient Low Rank Attention for Long-Context Inference in Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.23649