Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hara, Tomomasa, Kurita, Hiroto, Imaizumi, Masaaki, Inui, Kentaro, Yokoi, Sho
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914519979130880
author Hara, Tomomasa
Kurita, Hiroto
Imaizumi, Masaaki
Inui, Kentaro
Yokoi, Sho
author_facet Hara, Tomomasa
Kurita, Hiroto
Imaizumi, Masaaki
Inui, Kentaro
Yokoi, Sho
contents For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach. This paper examines whether mean pooling actually works well in real models. First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings. Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling. Then, using this metric, we empirically measure how often this collapse occurs in actual models and texts, and find that modern text encoders are robust to this collapse. In particular, contrastive fine-tuned text encoders tend to be less prone to the collapse than their pretrained backbone models. We also find that the robustness of these text encoders lies in the concentration of token embeddings within each text. In addition, we find that robustness to the collapse, as quantified by our proposed metric, correlates with downstream task performance. Overall, our findings offer a new perspective on why modern text encoders remain effective despite relying on seemingly coarse mean pooling.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27398
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
Hara, Tomomasa
Kurita, Hiroto
Imaizumi, Masaaki
Inui, Kentaro
Yokoi, Sho
Computation and Language
For constructing text embeddings, mean pooling, which averages token embeddings, is the standard approach. This paper examines whether mean pooling actually works well in real models. First, we note that mean pooling can collapse information beyond the first-order statistics of the token embeddings, such as second-order statistics that capture their spatial structure, potentially mapping distinct token embedding distributions to similar text embeddings. Motivated by this concern, we propose a simple metric to quantify such a collapse induced by mean pooling. Then, using this metric, we empirically measure how often this collapse occurs in actual models and texts, and find that modern text encoders are robust to this collapse. In particular, contrastive fine-tuned text encoders tend to be less prone to the collapse than their pretrained backbone models. We also find that the robustness of these text encoders lies in the concentration of token embeddings within each text. In addition, we find that robustness to the collapse, as quantified by our proposed metric, correlates with downstream task performance. Overall, our findings offer a new perspective on why modern text encoders remain effective despite relying on seemingly coarse mean pooling.
title Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
topic Computation and Language
url https://arxiv.org/abs/2604.27398