DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915196858007552 |
|---|---|
| author | Xu, Zhe Ye, Jiasheng Liu, Xiaoran Liu, Xiangyang Sun, Tianxiang Liu, Zhigeng Guo, Qipeng Li, Linlin Liu, Qun Huang, Xuanjing Qiu, Xipeng |
| author_facet | Xu, Zhe Ye, Jiasheng Liu, Xiaoran Liu, Xiangyang Sun, Tianxiang Liu, Zhigeng Guo, Qipeng Li, Linlin Liu, Qun Huang, Xuanjing Qiu, Xipeng |
| contents | Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset specifically designed for narrative reasoning within long contexts. We leverage detective novels, averaging over 100k tokens, to create a dataset containing 1200 human-annotated questions in both Chinese and English, each paired with corresponding reference reasoning steps. Furthermore, we introduce a step-wise reasoning metric, which enhances the evaluation of LLMs' reasoning processes. We validate our approach and evaluate the mainstream LLMs, including GPT-4, Claude, and LLaMA, revealing persistent long-context reasoning challenges and demonstrating their evidence-retrieval challenges. Our findings offer valuable insights into the study of long-context reasoning and lay the base for more rigorous evaluations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_02465 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels Xu, Zhe Ye, Jiasheng Liu, Xiaoran Liu, Xiangyang Sun, Tianxiang Liu, Zhigeng Guo, Qipeng Li, Linlin Liu, Qun Huang, Xuanjing Qiu, Xipeng Computation and Language Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset specifically designed for narrative reasoning within long contexts. We leverage detective novels, averaging over 100k tokens, to create a dataset containing 1200 human-annotated questions in both Chinese and English, each paired with corresponding reference reasoning steps. Furthermore, we introduce a step-wise reasoning metric, which enhances the evaluation of LLMs' reasoning processes. We validate our approach and evaluate the mainstream LLMs, including GPT-4, Claude, and LLaMA, revealing persistent long-context reasoning challenges and demonstrating their evidence-retrieval challenges. Our findings offer valuable insights into the study of long-context reasoning and lay the base for more rigorous evaluations. |
| title | DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2409.02465 |