FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Thakur, Nandan, Lin, Jimmy, Havens, Sam, Carbin, Michael, Khattab, Omar, Drozdov, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025)
by: Pradeep, Ronak, et al.
Published: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
by: Thakur, Nandan, et al.
Published: (2026)
by: Thakur, Nandan, et al.
Published: (2026)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
Backtracing: Retrieving the Cause of the Query
by: Wang, Rose E., et al.
Published: (2024)
by: Wang, Rose E., et al.
Published: (2024)
Building Russian Benchmark for Evaluation of Information Retrieval Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
by: Sharifymoghaddam, Sahel, et al.
Published: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
by: Chen, Zijian, et al.
Published: (2025)
by: Chen, Zijian, et al.
Published: (2025)
Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
by: Kim, To Eun, et al.
Published: (2024)
by: Kim, To Eun, et al.
Published: (2024)
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
by: Su, Hongjin, et al.
Published: (2024)
by: Su, Hongjin, et al.
Published: (2024)
Loops On Retrieval Augmented Generation (LoRAG)
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
AIANO: Enhancing Information Retrieval with AI-Augmented Annotation
by: Khattab, Sameh, et al.
Published: (2026)
by: Khattab, Sameh, et al.
Published: (2026)
DAPR: A Benchmark on Document-Aware Passage Retrieval
by: Wang, Kexin, et al.
Published: (2023)
by: Wang, Kexin, et al.
Published: (2023)
STELLA: Self-Reflective Terminology-Aware Framework for Building an Aerospace Information Retrieval Benchmark
by: Kim, Bongmin
Published: (2026)
by: Kim, Bongmin
Published: (2026)
Evaluating the Performance of LLMs on Technical Language Processing tasks
by: Kernycky, Andrew, et al.
Published: (2024)
by: Kernycky, Andrew, et al.
Published: (2024)
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
by: Acharya, Arkadeep, et al.
Published: (2024)
by: Acharya, Arkadeep, et al.
Published: (2024)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
A Systematic Study of Pseudo-Relevance Feedback with LLMs
by: Jedidi, Nour, et al.
Published: (2026)
by: Jedidi, Nour, et al.
Published: (2026)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
by: Yang, Jheng-Hong, et al.
Published: (2024)
by: Yang, Jheng-Hong, et al.
Published: (2024)
Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
OBLIQ-Bench: Exposing Overlooked Bottlenecks in Modern Retrievers with Latent and Implicit Queries
by: Tchuindjo, Diane, et al.
Published: (2026)
by: Tchuindjo, Diane, et al.
Published: (2026)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
by: Hsu, Tz-Huan, et al.
Published: (2026)
by: Hsu, Tz-Huan, et al.
Published: (2026)
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
by: Jagadeeshan, Manoj Balaji, et al.
Published: (2025)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
by: Opsahl-Ong, Krista, et al.
Published: (2026)
by: Opsahl-Ong, Krista, et al.
Published: (2026)
Similar Items
-
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026) -
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024) -
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025) -
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025) -
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)