DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Zhe, Ye, Jiasheng, Liu, Xiaoran, Liu, Xiangyang, Sun, Tianxiang, Liu, Zhigeng, Guo, Qipeng, Li, Linlin, Liu, Qun, Huang, Xuanjing, Qiu, Xipeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915196858007552
author Xu, Zhe
Ye, Jiasheng
Liu, Xiaoran
Liu, Xiangyang
Sun, Tianxiang
Liu, Zhigeng
Guo, Qipeng
Li, Linlin
Liu, Qun
Huang, Xuanjing
Qiu, Xipeng
author_facet Xu, Zhe
Ye, Jiasheng
Liu, Xiaoran
Liu, Xiangyang
Sun, Tianxiang
Liu, Zhigeng
Guo, Qipeng
Li, Linlin
Liu, Qun
Huang, Xuanjing
Qiu, Xipeng
contents Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset specifically designed for narrative reasoning within long contexts. We leverage detective novels, averaging over 100k tokens, to create a dataset containing 1200 human-annotated questions in both Chinese and English, each paired with corresponding reference reasoning steps. Furthermore, we introduce a step-wise reasoning metric, which enhances the evaluation of LLMs' reasoning processes. We validate our approach and evaluate the mainstream LLMs, including GPT-4, Claude, and LLaMA, revealing persistent long-context reasoning challenges and demonstrating their evidence-retrieval challenges. Our findings offer valuable insights into the study of long-context reasoning and lay the base for more rigorous evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02465
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
Xu, Zhe
Ye, Jiasheng
Liu, Xiaoran
Liu, Xiangyang
Sun, Tianxiang
Liu, Zhigeng
Guo, Qipeng
Li, Linlin
Liu, Qun
Huang, Xuanjing
Qiu, Xipeng
Computation and Language
Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset specifically designed for narrative reasoning within long contexts. We leverage detective novels, averaging over 100k tokens, to create a dataset containing 1200 human-annotated questions in both Chinese and English, each paired with corresponding reference reasoning steps. Furthermore, we introduce a step-wise reasoning metric, which enhances the evaluation of LLMs' reasoning processes. We validate our approach and evaluate the mainstream LLMs, including GPT-4, Claude, and LLaMA, revealing persistent long-context reasoning challenges and demonstrating their evidence-retrieval challenges. Our findings offer valuable insights into the study of long-context reasoning and lay the base for more rigorous evaluations.
title DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
topic Computation and Language
url https://arxiv.org/abs/2409.02465