FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Thakur, Nandan, Lin, Jimmy, Havens, Sam, Carbin, Michael, Khattab, Omar, Drozdov, Andrew
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908406196994048
author Thakur, Nandan
Lin, Jimmy
Havens, Sam
Carbin, Michael
Khattab, Omar
Drozdov, Andrew
author_facet Thakur, Nandan
Lin, Jimmy
Havens, Sam
Carbin, Michael
Khattab, Omar
Drozdov, Andrew
contents We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshStack conducts the following steps: (1) automatic corpus collection from code and technical documentation, (2) nugget generation from community-asked questions and answers, and (3) nugget-level support, retrieving documents using a fusion of retrieval techniques and hybrid architectures. We use FreshStack to build five datasets on fast-growing, recent, and niche topics to ensure the tasks are sufficiently challenging. On FreshStack, existing retrieval models, when applied out-of-the-box, significantly underperform oracle approaches on all five topics, denoting plenty of headroom to improve IR quality. In addition, we identify cases where rerankers do not improve first-stage retrieval accuracy (two out of five topics) and oracle context helps an LLM generator generate a high-quality RAG answer. We hope FreshStack will facilitate future work toward constructing realistic, scalable, and uncontaminated IR and RAG evaluation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13128
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
Thakur, Nandan
Lin, Jimmy
Havens, Sam
Carbin, Michael
Khattab, Omar
Drozdov, Andrew
Information Retrieval
Artificial Intelligence
Computation and Language
We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshStack conducts the following steps: (1) automatic corpus collection from code and technical documentation, (2) nugget generation from community-asked questions and answers, and (3) nugget-level support, retrieving documents using a fusion of retrieval techniques and hybrid architectures. We use FreshStack to build five datasets on fast-growing, recent, and niche topics to ensure the tasks are sufficiently challenging. On FreshStack, existing retrieval models, when applied out-of-the-box, significantly underperform oracle approaches on all five topics, denoting plenty of headroom to improve IR quality. In addition, we identify cases where rerankers do not improve first-stage retrieval accuracy (two out of five topics) and oracle context helps an LLM generator generate a high-quality RAG answer. We hope FreshStack will facilitate future work toward constructing realistic, scalable, and uncontaminated IR and RAG evaluation benchmarks.
title FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.13128