SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Biao, Chen, Lixin, Liu, Tong, Zheng, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911209357312000
author Zhang, Biao
Chen, Lixin
Liu, Tong
Zheng, Bo
author_facet Zhang, Biao
Chen, Lixin
Liu, Tong
Zheng, Bo
contents Large language models (LLMs) generate high-dimensional embeddings that capture rich semantic and syntactic information. However, high-dimensional embeddings exacerbate computational complexity and storage requirements, thereby hindering practical deployment. To address these challenges, we propose a novel training framework named Sequential Matryoshka Embedding Compression (SMEC). This framework introduces the Sequential Matryoshka Representation Learning(SMRL) method to mitigate gradient variance during training, the Adaptive Dimension Selection (ADS) module to reduce information degradation during dimension pruning, and the Selectable Cross-batch Memory (S-XBM) module to enhance unsupervised learning between high- and low-dimensional embeddings. Experiments on image, text, and multimodal datasets demonstrate that SMEC achieves significant dimensionality reduction while maintaining performance. For instance, on the BEIR dataset, our approach improves the performance of compressed LLM2Vec embeddings (256 dimensions) by 1.1 points and 2.7 points compared to the Matryoshka-Adaptor and Search-Adaptor models, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12474
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
Zhang, Biao
Chen, Lixin
Liu, Tong
Zheng, Bo
Computation and Language
Machine Learning
Large language models (LLMs) generate high-dimensional embeddings that capture rich semantic and syntactic information. However, high-dimensional embeddings exacerbate computational complexity and storage requirements, thereby hindering practical deployment. To address these challenges, we propose a novel training framework named Sequential Matryoshka Embedding Compression (SMEC). This framework introduces the Sequential Matryoshka Representation Learning(SMRL) method to mitigate gradient variance during training, the Adaptive Dimension Selection (ADS) module to reduce information degradation during dimension pruning, and the Selectable Cross-batch Memory (S-XBM) module to enhance unsupervised learning between high- and low-dimensional embeddings. Experiments on image, text, and multimodal datasets demonstrate that SMEC achieves significant dimensionality reduction while maintaining performance. For instance, on the BEIR dataset, our approach improves the performance of compressed LLM2Vec embeddings (256 dimensions) by 1.1 points and 2.7 points compared to the Matryoshka-Adaptor and Search-Adaptor models, respectively.
title SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.12474