To Case or Not to Case: An Empirical Study in Learned Sparse Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lionis, Emmanouil Georgios, Ju, Jia-Huei, Nalmpantis, Angelos, Thuis, Casper, MacAvaney, Sean, Yates, Andrew
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911396729454592
author Lionis, Emmanouil Georgios
Ju, Jia-Huei
Nalmpantis, Angelos
Thuis, Casper
MacAvaney, Sean
Yates, Andrew
author_facet Lionis, Emmanouil Georgios
Ju, Jia-Huei
Nalmpantis, Angelos
Thuis, Casper
MacAvaney, Sean
Yates, Andrew
contents Learned Sparse Retrieval (LSR) methods construct sparse lexical representations of queries and documents that can be efficiently searched using inverted indexes. Existing LSR approaches have relied almost exclusively on uncased backbone models, whose vocabularies exclude case-sensitive distinctions, thereby reducing vocabulary mismatch. However, the most recent state-of-the-art language models are only available in cased versions. Despite this shift, the impact of backbone model casing on LSR has not been studied, potentially posing a risk to the viability of the method going forward. To fill this gap, we systematically evaluate paired cased and uncased versions of the same backbone models across multiple datasets to assess their suitability for LSR. Our findings show that LSR models with cased backbone models by default perform substantially worse than their uncased counterparts; however, this gap can be eliminated by pre-processing the text to lowercase. Moreover, our token-level analysis reveals that, under lowercasing, cased models almost entirely suppress cased vocabulary items and behave effectively as uncased models, explaining their restored performance. This result broadens the applicability of recent cased models to the LSR setting and facilitates the integration of stronger backbone architectures into sparse retrieval. The complete code and implementation for this project are available at: https://github.com/lionisakis/Uncased-vs-cased-models-in-LSR
format Preprint
id arxiv_https___arxiv_org_abs_2601_17500
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle To Case or Not to Case: An Empirical Study in Learned Sparse Retrieval
Lionis, Emmanouil Georgios
Ju, Jia-Huei
Nalmpantis, Angelos
Thuis, Casper
MacAvaney, Sean
Yates, Andrew
Information Retrieval
Computation and Language
Learned Sparse Retrieval (LSR) methods construct sparse lexical representations of queries and documents that can be efficiently searched using inverted indexes. Existing LSR approaches have relied almost exclusively on uncased backbone models, whose vocabularies exclude case-sensitive distinctions, thereby reducing vocabulary mismatch. However, the most recent state-of-the-art language models are only available in cased versions. Despite this shift, the impact of backbone model casing on LSR has not been studied, potentially posing a risk to the viability of the method going forward. To fill this gap, we systematically evaluate paired cased and uncased versions of the same backbone models across multiple datasets to assess their suitability for LSR. Our findings show that LSR models with cased backbone models by default perform substantially worse than their uncased counterparts; however, this gap can be eliminated by pre-processing the text to lowercase. Moreover, our token-level analysis reveals that, under lowercasing, cased models almost entirely suppress cased vocabulary items and behave effectively as uncased models, explaining their restored performance. This result broadens the applicability of recent cased models to the LSR setting and facilitates the integration of stronger backbone architectures into sparse retrieval. The complete code and implementation for this project are available at: https://github.com/lionisakis/Uncased-vs-cased-models-in-LSR
title To Case or Not to Case: An Empirical Study in Learned Sparse Retrieval
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2601.17500