QKV Projections Require a Fraction of Their Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalaf, Malik, Shamshoum, Yara, Hodos, Nitzan, Sieradzki, Yuval, Schuster, Assaf
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910036174831616
author Khalaf, Malik
Shamshoum, Yara
Hodos, Nitzan
Sieradzki, Yuval
Schuster, Assaf
author_facet Khalaf, Malik
Shamshoum, Yara
Hodos, Nitzan
Sieradzki, Yuval
Schuster, Assaf
contents The Multi-Head Attention mechanism is central to LLM operation, and multiple works target its compute and memory efficiency during training. While most works focus on approximating the scaled dot product, the memory consumption of the linear projections that compute the $Q$, $K$, and $V$ tensors from the input $x$ is often overlooked. To address this, we propose Point-Approximate Matrix Multiplication (PAMM), a novel tensor compression technique that compresses the activations of the $Q,K,V$ projections in attention layers by a factor of up to $\times 512$, effectively erasing their memory footprint, while achieving similar or better final perplexity. PAMM is fully composable with efficient attention techniques such as FlashAttention, making it a practical and complementary method for memory-efficient LLM training.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QKV Projections Require a Fraction of Their Memory
Khalaf, Malik
Shamshoum, Yara
Hodos, Nitzan
Sieradzki, Yuval
Schuster, Assaf
Machine Learning
The Multi-Head Attention mechanism is central to LLM operation, and multiple works target its compute and memory efficiency during training. While most works focus on approximating the scaled dot product, the memory consumption of the linear projections that compute the $Q$, $K$, and $V$ tensors from the input $x$ is often overlooked. To address this, we propose Point-Approximate Matrix Multiplication (PAMM), a novel tensor compression technique that compresses the activations of the $Q,K,V$ projections in attention layers by a factor of up to $\times 512$, effectively erasing their memory footprint, while achieving similar or better final perplexity. PAMM is fully composable with efficient attention techniques such as FlashAttention, making it a practical and complementary method for memory-efficient LLM training.
title QKV Projections Require a Fraction of Their Memory
topic Machine Learning
url https://arxiv.org/abs/2506.02939