Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Zhouhong, Zhang, Lin, Chen, Jiangjie, Ye, Haoning, Zhu, Xiaoxuan, Li, Zihan, Ye, Zheyu, Gao, Yan, Hu, Yao, Xiao, Yanghua, Feng, Hongwei
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909143296638976
author Gu, Zhouhong
Zhang, Lin
Chen, Jiangjie
Ye, Haoning
Zhu, Xiaoxuan
Li, Zihan
Ye, Zheyu
Gao, Yan
Hu, Yao
Xiao, Yanghua
Feng, Hongwei
author_facet Gu, Zhouhong
Zhang, Lin
Chen, Jiangjie
Ye, Haoning
Zhu, Xiaoxuan
Li, Zihan
Ye, Zheyu
Gao, Yan
Hu, Yao
Xiao, Yanghua
Feng, Hongwei
contents Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of large language models~(LLMs), evaluating how these models identify key information and reason to solve questions becomes increasingly relevant. We introduces the DetectBench, a reading comprehension dataset designed to assess a model's ability to jointly ability in key information detection and multi-hop reasoning when facing complex and implicit information. The DetectBench comprises 3,928 questions, each paired with a paragraph averaging 190 tokens in length. To enhance model's detective skills, we propose the Detective Thinking Framework. These methods encourage models to identify all possible clues within the context before reasoning. Our experiments reveal that existing models perform poorly in both information detection and multi-hop reasoning. However, the Detective Thinking Framework approach alleviates this issue.
format Preprint
id arxiv_https___arxiv_org_abs_2307_05113
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models
Gu, Zhouhong
Zhang, Lin
Chen, Jiangjie
Ye, Haoning
Zhu, Xiaoxuan
Li, Zihan
Ye, Zheyu
Gao, Yan
Hu, Yao
Xiao, Yanghua
Feng, Hongwei
Computation and Language
Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of large language models~(LLMs), evaluating how these models identify key information and reason to solve questions becomes increasingly relevant. We introduces the DetectBench, a reading comprehension dataset designed to assess a model's ability to jointly ability in key information detection and multi-hop reasoning when facing complex and implicit information. The DetectBench comprises 3,928 questions, each paired with a paragraph averaging 190 tokens in length. To enhance model's detective skills, we propose the Detective Thinking Framework. These methods encourage models to identify all possible clues within the context before reasoning. Our experiments reveal that existing models perform poorly in both information detection and multi-hop reasoning. However, the Detective Thinking Framework approach alleviates this issue.
title Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2307.05113