LMK > CLS: Landmark Pooling for Dense Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Doshi, Meet, Trivedi, Aashka, Kumar, Vishwajeet, Awasthy, Parul, Li, Yulong, Sen, Jaydeep, Florian, Radu, Joshi, Sachindra
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911406254718976
author Doshi, Meet
Trivedi, Aashka
Kumar, Vishwajeet
Awasthy, Parul
Li, Yulong
Sen, Jaydeep
Florian, Radu
Joshi, Sachindra
author_facet Doshi, Meet
Trivedi, Aashka
Kumar, Vishwajeet
Awasthy, Parul
Li, Yulong
Sen, Jaydeep
Florian, Radu
Joshi, Sachindra
contents Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a variable-length token sequence to a single vector using a pooling operator, most commonly a special [CLS] token or mean pooling over token embeddings. In this paper, we identify systematic weaknesses of these pooling strategies: [CLS] tends to concentrate information toward the initial positions of the sequence and can under-represent distributed evidence, while mean pooling can dilute salient local signals, sometimes leading to worse short-context performance. To address these issues, we introduce Landmark (LMK) pooling, which partitions a sequence into chunks, inserts landmark tokens between chunks, and forms the final representation by mean-pooling the landmark token embeddings. This simple mechanism improves long-context extrapolation without sacrificing local salient features, at the cost of introducing a small number of special tokens. We empirically demonstrate that LMK pooling matches existing methods on short-context retrieval tasks and yields substantial improvements on long-context tasks, making it a practical and scalable alternative to existing pooling methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21525
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LMK > CLS: Landmark Pooling for Dense Embeddings
Doshi, Meet
Trivedi, Aashka
Kumar, Vishwajeet
Awasthy, Parul
Li, Yulong
Sen, Jaydeep
Florian, Radu
Joshi, Sachindra
Computation and Language
Information Retrieval
Representation learning is central to many downstream tasks such as search, clustering, classification, and reranking. State-of-the-art sequence encoders typically collapse a variable-length token sequence to a single vector using a pooling operator, most commonly a special [CLS] token or mean pooling over token embeddings. In this paper, we identify systematic weaknesses of these pooling strategies: [CLS] tends to concentrate information toward the initial positions of the sequence and can under-represent distributed evidence, while mean pooling can dilute salient local signals, sometimes leading to worse short-context performance. To address these issues, we introduce Landmark (LMK) pooling, which partitions a sequence into chunks, inserts landmark tokens between chunks, and forms the final representation by mean-pooling the landmark token embeddings. This simple mechanism improves long-context extrapolation without sacrificing local salient features, at the cost of introducing a small number of special tokens. We empirically demonstrate that LMK pooling matches existing methods on short-context retrieval tasks and yields substantial improvements on long-context tasks, making it a practical and scalable alternative to existing pooling methods.
title LMK > CLS: Landmark Pooling for Dense Embeddings
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2601.21525