Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sharifymoghaddam, Sahel, Ge, Yijun, Lin, Jimmy
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910277986942976
author Sharifymoghaddam, Sahel
Ge, Yijun
Lin, Jimmy
author_facet Sharifymoghaddam, Sahel
Ge, Yijun
Lin, Jimmy
contents Retrieval benchmarks for large language models (LLMs) should reflect the long, reasoning-intensive queries typical of retrieval-augmented generation (RAG). We present a systematic study of BRIGHT, a reasoning-focused retrieval benchmark, along with strong, reproducible reference methods integrated into Anserini, Pyserini, and RankLLM. We evaluate lexical, sparse, dense, and fusion-based retrievers, as well as LLM rerankers, under long-query settings. In reproducing BRIGHT's lexical baseline, we identify a key under-documented detail: query-side BM25 (BM25Q), which applies BM25 weighting to the query itself. On long, multi-sentence queries, BM25Q consistently outperforms standard BM25, making it the strongest lexical baseline for reasoning-oriented retrieval. We further audit the BRIGHT corpus, uncovering data quality issues that impact evaluation, and offer mitigation. Finally, we study the generalizability of BM25Q across five additional benchmarks, finding its gains largely specific to BRIGHT, while fusion with standard BM25 provides the most consistent improvements across datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
Sharifymoghaddam, Sahel
Ge, Yijun
Lin, Jimmy
Information Retrieval
Retrieval benchmarks for large language models (LLMs) should reflect the long, reasoning-intensive queries typical of retrieval-augmented generation (RAG). We present a systematic study of BRIGHT, a reasoning-focused retrieval benchmark, along with strong, reproducible reference methods integrated into Anserini, Pyserini, and RankLLM. We evaluate lexical, sparse, dense, and fusion-based retrievers, as well as LLM rerankers, under long-query settings. In reproducing BRIGHT's lexical baseline, we identify a key under-documented detail: query-side BM25 (BM25Q), which applies BM25 weighting to the query itself. On long, multi-sentence queries, BM25Q consistently outperforms standard BM25, making it the strongest lexical baseline for reasoning-oriented retrieval. We further audit the BRIGHT corpus, uncovering data quality issues that impact evaluation, and offer mitigation. Finally, we study the generalizability of BM25Q across five additional benchmarks, finding its gains largely specific to BRIGHT, while fusion with standard BM25 provides the most consistent improvements across datasets.
title Lighting the Way for BRIGHT: Reproducible Baselines with Anserini, Pyserini, and RankLLM
topic Information Retrieval
url https://arxiv.org/abs/2509.02558