Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Kuissi, Nathan, Subrahmanyan, Suraj, Thakur, Nandan, Lin, Jimmy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
por: Thakur, Nandan, et al.
Publicado: (2025)
por: Thakur, Nandan, et al.
Publicado: (2025)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
por: Thakur, Nandan, et al.
Publicado: (2025)
por: Thakur, Nandan, et al.
Publicado: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
por: Thakur, Nandan, et al.
Publicado: (2026)
por: Thakur, Nandan, et al.
Publicado: (2026)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
por: Thakur, Nandan, et al.
Publicado: (2023)
por: Thakur, Nandan, et al.
Publicado: (2023)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
por: Thakur, Nandan, et al.
Publicado: (2025)
por: Thakur, Nandan, et al.
Publicado: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
por: Pradeep, Ronak, et al.
Publicado: (2024)
por: Pradeep, Ronak, et al.
Publicado: (2024)
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
por: Hsu, Tz-Huan, et al.
Publicado: (2026)
por: Hsu, Tz-Huan, et al.
Publicado: (2026)
A Survey on Retrieval-Augmented Text Generation for Large Language Models
por: Huang, Yizheng, et al.
Publicado: (2024)
por: Huang, Yizheng, et al.
Publicado: (2024)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
por: Pradeep, Ronak, et al.
Publicado: (2025)
por: Pradeep, Ronak, et al.
Publicado: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
por: Pradeep, Ronak, et al.
Publicado: (2024)
por: Pradeep, Ronak, et al.
Publicado: (2024)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
por: Killingback, Julian, et al.
Publicado: (2025)
por: Killingback, Julian, et al.
Publicado: (2025)
Benchmarking Retrieval-Augmented Generation for Chemistry
por: Zhong, Xianrui, et al.
Publicado: (2025)
por: Zhong, Xianrui, et al.
Publicado: (2025)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
por: Kim, Eric Y., et al.
Publicado: (2026)
por: Kim, Eric Y., et al.
Publicado: (2026)
Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
por: Shi, Zhengliang, et al.
Publicado: (2025)
por: Shi, Zhengliang, et al.
Publicado: (2025)
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
por: Huang, Yizheng, et al.
Publicado: (2024)
por: Huang, Yizheng, et al.
Publicado: (2024)
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges
por: Wang, Jiajia, et al.
Publicado: (2024)
por: Wang, Jiajia, et al.
Publicado: (2024)
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
por: Su, Hongjin, et al.
Publicado: (2024)
por: Su, Hongjin, et al.
Publicado: (2024)
SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions
por: Mishra, Saroj, et al.
Publicado: (2026)
por: Mishra, Saroj, et al.
Publicado: (2026)
BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
por: Wojtasik, Konrad, et al.
Publicado: (2023)
por: Wojtasik, Konrad, et al.
Publicado: (2023)
Reliable Evaluation Protocol for Low-Precision Retrieval
por: Yang, Kisu, et al.
Publicado: (2025)
por: Yang, Kisu, et al.
Publicado: (2025)
Evaluation of retrieval-based QA on QUEST-LOFT
por: Scales, Nathan, et al.
Publicado: (2025)
por: Scales, Nathan, et al.
Publicado: (2025)
URAG: A Benchmark for Uncertainty Quantification in Retrieval-Augmented Large Language Models
por: Nguyen, Vinh, et al.
Publicado: (2026)
por: Nguyen, Vinh, et al.
Publicado: (2026)
Overview of the TREC 2021 deep learning track
por: Craswell, Nick, et al.
Publicado: (2025)
por: Craswell, Nick, et al.
Publicado: (2025)
Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study
por: Yu, Shuo, et al.
Publicado: (2024)
por: Yu, Shuo, et al.
Publicado: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
por: Jin, Zhuoran, et al.
Publicado: (2024)
por: Jin, Zhuoran, et al.
Publicado: (2024)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
por: Choi, Chanyeol, et al.
Publicado: (2025)
por: Choi, Chanyeol, et al.
Publicado: (2025)
Towards Personalized Deep Research: Benchmarks and Evaluations
por: Liang, Yuan, et al.
Publicado: (2025)
por: Liang, Yuan, et al.
Publicado: (2025)
TARGET: Benchmarking Table Retrieval for Generative Tasks
por: Ji, Xingyu, et al.
Publicado: (2025)
por: Ji, Xingyu, et al.
Publicado: (2025)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
por: Chi, Yizhou, et al.
Publicado: (2024)
por: Chi, Yizhou, et al.
Publicado: (2024)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
por: Saad-Falcon, Jon, et al.
Publicado: (2023)
por: Saad-Falcon, Jon, et al.
Publicado: (2023)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
por: Ezerceli, Özay, et al.
Publicado: (2025)
por: Ezerceli, Özay, et al.
Publicado: (2025)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
por: Khanghah, Kiarash Naghavi, et al.
Publicado: (2026)
por: Khanghah, Kiarash Naghavi, et al.
Publicado: (2026)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
por: Merola, Carlo, et al.
Publicado: (2025)
por: Merola, Carlo, et al.
Publicado: (2025)
MMTEB: Massive Multilingual Text Embedding Benchmark
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
Optimizing Multi-Hop Document Retrieval Through Intermediate Representations
por: Lin, Jiaen, et al.
Publicado: (2025)
por: Lin, Jiaen, et al.
Publicado: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
por: Ngo, Nghia Trung, et al.
Publicado: (2024)
por: Ngo, Nghia Trung, et al.
Publicado: (2024)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
por: Balog, Krisztian, et al.
Publicado: (2025)
por: Balog, Krisztian, et al.
Publicado: (2025)
Generative Retrieval and Alignment Model: A New Paradigm for E-commerce Retrieval
por: Pang, Ming, et al.
Publicado: (2025)
por: Pang, Ming, et al.
Publicado: (2025)
Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
por: Tang, Qiaoyu, et al.
Publicado: (2024)
por: Tang, Qiaoyu, et al.
Publicado: (2024)
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
por: Su, Weihang, et al.
Publicado: (2025)
por: Su, Weihang, et al.
Publicado: (2025)
Ejemplares similares
-
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
por: Thakur, Nandan, et al.
Publicado: (2025) -
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
por: Thakur, Nandan, et al.
Publicado: (2025) -
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
por: Thakur, Nandan, et al.
Publicado: (2026) -
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
por: Thakur, Nandan, et al.
Publicado: (2023) -
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
por: Thakur, Nandan, et al.
Publicado: (2025)