NeoQA: Evidence-based Question Answering with Generated News Events

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Glockner, Max, Jiang, Xiang, Ribeiro, Leonardo F. R., Gurevych, Iryna, Dreyer, Markus
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916728756240384
author Glockner, Max
Jiang, Xiang
Ribeiro, Leonardo F. R.
Gurevych, Iryna
Dreyer, Markus
author_facet Glockner, Max
Jiang, Xiang
Ribeiro, Leonardo F. R.
Gurevych, Iryna
Dreyer, Markus
contents Evaluating Retrieval-Augmented Generation (RAG) in large language models (LLMs) is challenging because benchmarks can quickly become stale. Questions initially requiring retrieval may become answerable from pretraining knowledge as newer models incorporate more recent information during pretraining, making it difficult to distinguish evidence-based reasoning from recall. We introduce NeoQA (News Events for Out-of-training Question Answering), a benchmark designed to address this issue. To construct NeoQA, we generated timelines and knowledge bases of fictional news events and entities along with news articles and Q\&A pairs to prevent LLMs from leveraging pretraining knowledge, ensuring that no prior evidence exists in their training data. We propose our dataset as a new platform for evaluating evidence-based question answering, as it requires LLMs to generate responses exclusively from retrieved evidence and only when sufficient evidence is available. NeoQA enables controlled evaluation across various evidence scenarios, including cases with missing or misleading details. Our findings indicate that LLMs struggle to distinguish subtle mismatches between questions and evidence, and suffer from short-cut reasoning when key information required to answer a question is missing from the evidence, underscoring key limitations in evidence-based reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NeoQA: Evidence-based Question Answering with Generated News Events
Glockner, Max
Jiang, Xiang
Ribeiro, Leonardo F. R.
Gurevych, Iryna
Dreyer, Markus
Computation and Language
Evaluating Retrieval-Augmented Generation (RAG) in large language models (LLMs) is challenging because benchmarks can quickly become stale. Questions initially requiring retrieval may become answerable from pretraining knowledge as newer models incorporate more recent information during pretraining, making it difficult to distinguish evidence-based reasoning from recall. We introduce NeoQA (News Events for Out-of-training Question Answering), a benchmark designed to address this issue. To construct NeoQA, we generated timelines and knowledge bases of fictional news events and entities along with news articles and Q\&A pairs to prevent LLMs from leveraging pretraining knowledge, ensuring that no prior evidence exists in their training data. We propose our dataset as a new platform for evaluating evidence-based question answering, as it requires LLMs to generate responses exclusively from retrieved evidence and only when sufficient evidence is available. NeoQA enables controlled evaluation across various evidence scenarios, including cases with missing or misleading details. Our findings indicate that LLMs struggle to distinguish subtle mismatches between questions and evidence, and suffer from short-cut reasoning when key information required to answer a question is missing from the evidence, underscoring key limitations in evidence-based reasoning.
title NeoQA: Evidence-based Question Answering with Generated News Events
topic Computation and Language
url https://arxiv.org/abs/2505.05949