BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsirigotis, Christos, Adlakha, Vaibhav, Monteiro, Joao, Courville, Aaron, Taslakian, Perouz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909730879832064
author Tsirigotis, Christos
Adlakha, Vaibhav
Monteiro, Joao
Courville, Aaron
Taslakian, Perouz
author_facet Tsirigotis, Christos
Adlakha, Vaibhav
Monteiro, Joao
Courville, Aaron
Taslakian, Perouz
contents Neural sentence embedding models for dense retrieval typically rely on binary relevance labels, treating query-document pairs as either relevant or irrelevant. However, real-world relevance often exists on a continuum, and recent advances in large language models (LLMs) have made it feasible to scale the generation of fine-grained graded relevance labels. In this work, we propose BiXSE, a simple and effective pointwise training method that optimizes binary cross-entropy (BCE) over LLM-generated graded relevance scores. BiXSE interprets these scores as probabilistic targets, enabling granular supervision from a single labeled query-document pair per query. Unlike pairwise or listwise losses that require multiple annotated comparisons per query, BiXSE achieves strong performance with reduced annotation and compute costs by leveraging in-batch negatives. Extensive experiments across sentence embedding (MMTEB) and retrieval benchmarks (BEIR, TREC-DL) show that BiXSE consistently outperforms softmax-based contrastive learning (InfoNCE), and matches or exceeds strong pairwise ranking baselines when trained on LLM-supervised data. BiXSE offers a robust, scalable alternative for training dense retrieval models as graded relevance supervision becomes increasingly accessible.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
Tsirigotis, Christos
Adlakha, Vaibhav
Monteiro, Joao
Courville, Aaron
Taslakian, Perouz
Information Retrieval
Artificial Intelligence
Machine Learning
Neural sentence embedding models for dense retrieval typically rely on binary relevance labels, treating query-document pairs as either relevant or irrelevant. However, real-world relevance often exists on a continuum, and recent advances in large language models (LLMs) have made it feasible to scale the generation of fine-grained graded relevance labels. In this work, we propose BiXSE, a simple and effective pointwise training method that optimizes binary cross-entropy (BCE) over LLM-generated graded relevance scores. BiXSE interprets these scores as probabilistic targets, enabling granular supervision from a single labeled query-document pair per query. Unlike pairwise or listwise losses that require multiple annotated comparisons per query, BiXSE achieves strong performance with reduced annotation and compute costs by leveraging in-batch negatives. Extensive experiments across sentence embedding (MMTEB) and retrieval benchmarks (BEIR, TREC-DL) show that BiXSE consistently outperforms softmax-based contrastive learning (InfoNCE), and matches or exceeds strong pairwise ranking baselines when trained on LLM-supervised data. BiXSE offers a robust, scalable alternative for training dense retrieval models as graded relevance supervision becomes increasingly accessible.
title BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.06781