NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gupta, Abhay, Lu, Michael, Zhu, Kevin, O'Brien, Sean, Sharma, Vasu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909655526014976
author Gupta, Abhay
Lu, Michael
Zhu, Kevin
O'Brien, Sean
Sharma, Vasu
author_facet Gupta, Abhay
Lu, Michael
Zhu, Kevin
O'Brien, Sean
Sharma, Vasu
contents Current large language models (LLMs) struggle to answer questions that span tens of thousands of tokens, especially when multi-hop reasoning is involved. While prior benchmarks explore long-context comprehension or multi-hop reasoning in isolation, none jointly vary context length and reasoning depth in natural narrative settings. We introduce NovelHopQA, the first benchmark to evaluate 1-4 hop QA over 64k-128k-token excerpts from 83 full-length public-domain novels. A keyword-guided pipeline builds hop-separated chains grounded in coherent storylines. We evaluate seven state-of-the-art models and apply oracle-context filtering to ensure all questions are genuinely answerable. Human annotators validate both alignment and hop depth. We additionally present retrieval-augmented generation (RAG) evaluations to test model performance when only selected passages are provided instead of the full context. We noticed consistent accuracy drops with increased hops and context length increase, even for frontier models-revealing that sheer scale does not guarantee robust reasoning. Failure-mode analysis highlights common breakdowns such as missed final-hop integration and long-range drift. NovelHopQA offers a controlled diagnostic setting to test multi-hop reasoning at scale. All code and datasets are available at https://novelhopqa.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02000
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
Gupta, Abhay
Lu, Michael
Zhu, Kevin
O'Brien, Sean
Sharma, Vasu
Computation and Language
Current large language models (LLMs) struggle to answer questions that span tens of thousands of tokens, especially when multi-hop reasoning is involved. While prior benchmarks explore long-context comprehension or multi-hop reasoning in isolation, none jointly vary context length and reasoning depth in natural narrative settings. We introduce NovelHopQA, the first benchmark to evaluate 1-4 hop QA over 64k-128k-token excerpts from 83 full-length public-domain novels. A keyword-guided pipeline builds hop-separated chains grounded in coherent storylines. We evaluate seven state-of-the-art models and apply oracle-context filtering to ensure all questions are genuinely answerable. Human annotators validate both alignment and hop depth. We additionally present retrieval-augmented generation (RAG) evaluations to test model performance when only selected passages are provided instead of the full context. We noticed consistent accuracy drops with increased hops and context length increase, even for frontier models-revealing that sheer scale does not guarantee robust reasoning. Failure-mode analysis highlights common breakdowns such as missed final-hop integration and long-range drift. NovelHopQA offers a controlled diagnostic setting to test multi-hop reasoning at scale. All code and datasets are available at https://novelhopqa.github.io.
title NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
topic Computation and Language
url https://arxiv.org/abs/2506.02000