The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kasalický, Petr, Spišák, Martin, Vančura, Vojtěch, Bohuněk, Daniel, Alves, Rodrigo, Kordík, Pavel
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913842412388352
author Kasalický, Petr
Spišák, Martin
Vančura, Vojtěch
Bohuněk, Daniel
Alves, Rodrigo
Kordík, Pavel
author_facet Kasalický, Petr
Spišák, Martin
Vančura, Vojtěch
Bohuněk, Daniel
Alves, Rodrigo
Kordík, Pavel
contents Industry-scale recommender systems face a core challenge: representing entities with high cardinality, such as users or items, using dense embeddings that must be accessible during both training and inference. However, as embedding sizes grow, memory constraints make storage and access increasingly difficult. We describe a lightweight, learnable embedding compression technique that projects dense embeddings into a high-dimensional, sparsely activated space. Designed for retrieval tasks, our method reduces memory requirements while preserving retrieval performance, enabling scalable deployment under strict resource constraints. Our results demonstrate that leveraging sparsity is a promising approach for improving the efficiency of large-scale recommenders. We release our code at https://github.com/recombee/CompresSAE.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems
Kasalický, Petr
Spišák, Martin
Vančura, Vojtěch
Bohuněk, Daniel
Alves, Rodrigo
Kordík, Pavel
Information Retrieval
Machine Learning
Industry-scale recommender systems face a core challenge: representing entities with high cardinality, such as users or items, using dense embeddings that must be accessible during both training and inference. However, as embedding sizes grow, memory constraints make storage and access increasingly difficult. We describe a lightweight, learnable embedding compression technique that projects dense embeddings into a high-dimensional, sparsely activated space. Designed for retrieval tasks, our method reduces memory requirements while preserving retrieval performance, enabling scalable deployment under strict resource constraints. Our results demonstrate that leveraging sparsity is a promising approach for improving the efficiency of large-scale recommenders. We release our code at https://github.com/recombee/CompresSAE.
title The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2505.11388