Clustering-driven Memory Compression for On-device Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bohdal, Ondrej, Saha, Pramit, Michieli, Umberto, Ozay, Mete, Ceritli, Taha
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914277638537216
author Bohdal, Ondrej
Saha, Pramit
Michieli, Umberto
Ozay, Mete
Ceritli, Taha
author_facet Bohdal, Ondrej
Saha, Pramit
Michieli, Umberto
Ozay, Mete
Ceritli, Taha
contents Large language models (LLMs) often rely on user-specific memories distilled from past interactions to enable personalized generation. A common practice is to concatenate these memories with the input prompt, but this approach quickly exhausts the limited context available in on-device LLMs. Compressing memories by averaging can mitigate context growth, yet it frequently harms performance due to semantic conflicts across heterogeneous memories. In this work, we introduce a clustering-based memory compression strategy that balances context efficiency and personalization quality. Our method groups memories by similarity and merges them within clusters prior to concatenation, thereby preserving coherence while reducing redundancy. Experiments demonstrate that our approach substantially lowers the number of memory tokens while outperforming baseline strategies such as naive averaging or direct concatenation. Furthermore, for a fixed context budget, clustering-driven merging yields more compact memory representations and consistently enhances generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17443
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clustering-driven Memory Compression for On-device Large Language Models
Bohdal, Ondrej
Saha, Pramit
Michieli, Umberto
Ozay, Mete
Ceritli, Taha
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) often rely on user-specific memories distilled from past interactions to enable personalized generation. A common practice is to concatenate these memories with the input prompt, but this approach quickly exhausts the limited context available in on-device LLMs. Compressing memories by averaging can mitigate context growth, yet it frequently harms performance due to semantic conflicts across heterogeneous memories. In this work, we introduce a clustering-based memory compression strategy that balances context efficiency and personalization quality. Our method groups memories by similarity and merges them within clusters prior to concatenation, thereby preserving coherence while reducing redundancy. Experiments demonstrate that our approach substantially lowers the number of memory tokens while outperforming baseline strategies such as naive averaging or direct concatenation. Furthermore, for a fixed context budget, clustering-driven merging yields more compact memory representations and consistently enhances generation quality.
title Clustering-driven Memory Compression for On-device Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.17443