Evaluation of retrieval-based QA on QUEST-LOFT
Fuente:
arXiv
Saved in:
| Main Authors: | Scales, Nathan, Schärli, Nathanael, Bousquet, Olivier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the impact of retrieved content representations in RAG Pipelines
by: Ross, Jonathan J, et al.
Published: (2026)
by: Ross, Jonathan J, et al.
Published: (2026)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
by: Islamaj, Rezarta, et al.
Published: (2026)
by: Islamaj, Rezarta, et al.
Published: (2026)
MapQA: Open-domain Geospatial Question Answering on Map Data
by: Li, Zekun, et al.
Published: (2025)
by: Li, Zekun, et al.
Published: (2025)
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
by: Alushi, Klejda, et al.
Published: (2026)
by: Alushi, Klejda, et al.
Published: (2026)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
by: Opsahl-Ong, Krista, et al.
Published: (2026)
by: Opsahl-Ong, Krista, et al.
Published: (2026)
Fine-tune the Entire RAG Architecture (including DPR retriever) for Question-Answering
by: Siriwardhana, Shamane, et al.
Published: (2021)
by: Siriwardhana, Shamane, et al.
Published: (2021)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA
by: Ou, Justice, et al.
Published: (2025)
by: Ou, Justice, et al.
Published: (2025)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
by: Baumgärtner, Tim, et al.
Published: (2025)
by: Baumgärtner, Tim, et al.
Published: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
by: Kuissi, Nathan, et al.
Published: (2026)
by: Kuissi, Nathan, et al.
Published: (2026)
Horizon Scans can be accelerated using novel information retrieval and artificial intelligence tools
by: Schmidt, Lena, et al.
Published: (2025)
by: Schmidt, Lena, et al.
Published: (2025)
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-$k$
by: Taguchi, Chihiro, et al.
Published: (2025)
by: Taguchi, Chihiro, et al.
Published: (2025)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
LLM2IR: simple unsupervised contrastive learning makes long-context LLM great retriever
by: Yang, Xiaocong
Published: (2025)
by: Yang, Xiaocong
Published: (2025)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
by: Yang, Hang, et al.
Published: (2024)
by: Yang, Hang, et al.
Published: (2024)
WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2026)
by: Dinzinger, Michael, et al.
Published: (2026)
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV
by: Stinard, Alex
Published: (2026)
by: Stinard, Alex
Published: (2026)
Unlocking Electronic Health Records: A Hybrid Graph RAG Approach to Safe Clinical AI for Patient QA
by: Thio, Samuel, et al.
Published: (2025)
by: Thio, Samuel, et al.
Published: (2025)
HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
Med-CoDE: Medical Critique based Disagreement Evaluation Framework
by: Gupta, Mohit, et al.
Published: (2025)
by: Gupta, Mohit, et al.
Published: (2025)
Complex QA and language models hybrid architectures, Survey
by: Daull, Xavier, et al.
Published: (2023)
by: Daull, Xavier, et al.
Published: (2023)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
by: Zhang, Rongyang, et al.
Published: (2025)
by: Zhang, Rongyang, et al.
Published: (2025)
Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books
by: Papoudakis, Argyrios, et al.
Published: (2026)
by: Papoudakis, Argyrios, et al.
Published: (2026)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
by: Sun, Yiqun, et al.
Published: (2026)
by: Sun, Yiqun, et al.
Published: (2026)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
RCSB PDB AI Help Desk: retrieval-augmented generation for protein structure deposition support
by: Chithari, Vivek Reddy, et al.
Published: (2026)
by: Chithari, Vivek Reddy, et al.
Published: (2026)
Towards Personalized Deep Research: Benchmarks and Evaluations
by: Liang, Yuan, et al.
Published: (2025)
by: Liang, Yuan, et al.
Published: (2025)
RAVine: Reality-Aligned Evaluation for Agentic Search
by: Xu, Yilong, et al.
Published: (2025)
by: Xu, Yilong, et al.
Published: (2025)
Reliable Evaluation Protocol for Low-Precision Retrieval
by: Yang, Kisu, et al.
Published: (2025)
by: Yang, Kisu, et al.
Published: (2025)
Reindex-Then-Adapt: Improving Large Language Models for Conversational Recommendation
by: He, Zhankui, et al.
Published: (2024)
by: He, Zhankui, et al.
Published: (2024)
EncouRAGe: Evaluating RAG Local, Fast, and Reliable
by: Strich, Jan, et al.
Published: (2025)
by: Strich, Jan, et al.
Published: (2025)
Auto-ARGUE: LLM-Based Report Generation Evaluation
by: Walden, William, et al.
Published: (2025)
by: Walden, William, et al.
Published: (2025)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
by: Aggarwal, Shashank, et al.
Published: (2026)
by: Aggarwal, Shashank, et al.
Published: (2026)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
by: Xu, Yumo, et al.
Published: (2025)
by: Xu, Yumo, et al.
Published: (2025)
Efficient Evaluation of Large Language Models via Collaborative Filtering
by: Zhong, Xu-Xiang, et al.
Published: (2025)
by: Zhong, Xu-Xiang, et al.
Published: (2025)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
by: Hu, Xuming, et al.
Published: (2024)
by: Hu, Xuming, et al.
Published: (2024)
Evaluating the External and Parametric Knowledge Fusion of Large Language Models
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Similar Items
-
On the impact of retrieved content representations in RAG Pipelines
by: Ross, Jonathan J, et al.
Published: (2026) -
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
by: Islamaj, Rezarta, et al.
Published: (2026) -
MapQA: Open-domain Geospatial Question Answering on Map Data
by: Li, Zekun, et al.
Published: (2025) -
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
by: Alushi, Klejda, et al.
Published: (2026) -
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
by: Opsahl-Ong, Krista, et al.
Published: (2026)