Lattice: Learning to Efficiently Compress the Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karami, Mahdi, Pascanu, Razvan, Mirrokni, Vahab
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908739038085120
author Karami, Mahdi
Pascanu, Razvan
Mirrokni, Vahab
author_facet Karami, Mahdi
Pascanu, Razvan
Mirrokni, Vahab
contents Attention mechanisms have revolutionized sequence learning but suffer from quadratic computational complexity. This paper introduces \model, a novel recurrent neural network (RNN) mechanism that leverages the inherent low-rank structure of K-V matrices to efficiently compress the cache into a fixed number of memory slots, achieving sub-quadratic complexity. We formulate this compression as an online optimization problem and derive a dynamic memory update rule based on a single gradient descent step. The resulting recurrence features a state- and input-dependent gating mechanism, offering an interpretable memory update process. The core innovation is the orthogonal update: each memory slot is updated exclusively with information orthogonal to its current state, hence incorporating only novel, non-redundant data to minimize interference with previously stored information. We derive an efficient computation for this orthogonal update rule and further approximate it with chunk-wise parallelization to ensure training scalability. Empirically, Lattice outperforms strong baselines on language modeling and associative recall tasks across diverse context lengths and model sizes, achieving superior memory efficiency with significantly reduced memory sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lattice: Learning to Efficiently Compress the Memory
Karami, Mahdi
Pascanu, Razvan
Mirrokni, Vahab
Machine Learning
Artificial Intelligence
Attention mechanisms have revolutionized sequence learning but suffer from quadratic computational complexity. This paper introduces \model, a novel recurrent neural network (RNN) mechanism that leverages the inherent low-rank structure of K-V matrices to efficiently compress the cache into a fixed number of memory slots, achieving sub-quadratic complexity. We formulate this compression as an online optimization problem and derive a dynamic memory update rule based on a single gradient descent step. The resulting recurrence features a state- and input-dependent gating mechanism, offering an interpretable memory update process. The core innovation is the orthogonal update: each memory slot is updated exclusively with information orthogonal to its current state, hence incorporating only novel, non-redundant data to minimize interference with previously stored information. We derive an efficient computation for this orthogonal update rule and further approximate it with chunk-wise parallelization to ensure training scalability. Empirically, Lattice outperforms strong baselines on language modeling and associative recall tasks across diverse context lengths and model sizes, achieving superior memory efficiency with significantly reduced memory sizes.
title Lattice: Learning to Efficiently Compress the Memory
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.05646