PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yutao, Liu, Xiao, Feng, Yunhao, Ding, Jiale, Ma, Xingjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912670158946304
author Wu, Yutao
Liu, Xiao
Feng, Yunhao
Ding, Jiale
Ma, Xingjun
author_facet Wu, Yutao
Liu, Xiao
Feng, Yunhao
Ding, Jiale
Ma, Xingjun
contents Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a benchmark that systematically evaluates LLMs across four key research tasks: citation retrieval, content extraction, paper discovery, and claim verification. We evaluate GPT-4o, GPT-5, and Gemini-2.5-Flash under realistic usage conditions-via web interfaces where search operations are opaque to the user. Through controlled experiments, we find consistent reliability failures: citation retrieval fails in 48-98% of multi-reference queries, section-specific content extraction fails in 72-91% of cases, and topical paper discovery yields F1 scores below 0.32, missing over 60% of relevant literature. Further human analysis attributes these failures to the uncontrolled expansion of retrieved context and the tendency of LLMs to prioritize semantically relevant text over task instructions. Across basic tasks, the LLMs display distinct failure behaviors: ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers. To address these issues, we develop lightweight reliability classifiers trained on PaperAsk data to identify unreliable outputs. PaperAsk provides a reproducible and diagnostic framework for advancing the reliability evaluation of LLM-based scholarly assistance systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
Wu, Yutao
Liu, Xiao
Feng, Yunhao
Ding, Jiale
Ma, Xingjun
Information Retrieval
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a benchmark that systematically evaluates LLMs across four key research tasks: citation retrieval, content extraction, paper discovery, and claim verification. We evaluate GPT-4o, GPT-5, and Gemini-2.5-Flash under realistic usage conditions-via web interfaces where search operations are opaque to the user. Through controlled experiments, we find consistent reliability failures: citation retrieval fails in 48-98% of multi-reference queries, section-specific content extraction fails in 72-91% of cases, and topical paper discovery yields F1 scores below 0.32, missing over 60% of relevant literature. Further human analysis attributes these failures to the uncontrolled expansion of retrieved context and the tendency of LLMs to prioritize semantically relevant text over task instructions. Across basic tasks, the LLMs display distinct failure behaviors: ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers. To address these issues, we develop lightweight reliability classifiers trained on PaperAsk data to identify unreliable outputs. PaperAsk provides a reproducible and diagnostic framework for advancing the reliability evaluation of LLM-based scholarly assistance systems.
title PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.22242