PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Yanjun, Wei, Tianxin, Zou, Jiaru, Ning, Xuying, Bei, Yuanchen, Chen, Lingjie, Rana, Simmi, Yang, Wendy H., Tong, Hanghang, He, Jingrui
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908996756045824
author Zhao, Yanjun
Wei, Tianxin
Zou, Jiaru
Ning, Xuying
Bei, Yuanchen
Chen, Lingjie
Rana, Simmi
Yang, Wendy H.
Tong, Hanghang
He, Jingrui
author_facet Zhao, Yanjun
Wei, Tianxin
Zou, Jiaru
Ning, Xuying
Bei, Yuanchen
Chen, Lingjie
Rana, Simmi
Yang, Wendy H.
Tong, Hanghang
He, Jingrui
contents Understanding scientific papers requires more than answering isolated questions or summarizing content. It involves an integrated reasoning process that grounds textual and visual information, interprets experimental evidence, synthesizes information across sources, and critically evaluates scientific claims. However, existing benchmarks typically assess these abilities in isolation, making it difficult to evaluate scientific paper understanding as a unified set of interacting cognitive abilities. In this work, we introduce PaperMind, a benchmark designed to evaluate integrated and agent-oriented scientific reasoning over research papers. PaperMind is constructed from real scientific papers across seven domains, including agriculture, biology, chemistry, computer science, medicine, physics, and economics. It comprises four complementary task families that collectively operationalize distinct cognitive facets of scientific paper reasoning, including multimodal grounding, experimental interpretation, cross-source evidence reasoning, and critical assessment. By analyzing model behavior across multiple tasks, PaperMind enables a diagnostic evaluation of integrated scientific reasoning behaviors that are difficult to assess through isolated task evaluations. Extensive experiments on both opensource and closed-source multimodal LLMs reveal consistent performance gaps across tasks, highlighting persistent challenges in integrated scientific reasoning and critique. Our benchmark and dataset are available at https:// github.com/Yanjun-Zhao/PaperMind.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21304
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
Zhao, Yanjun
Wei, Tianxin
Zou, Jiaru
Ning, Xuying
Bei, Yuanchen
Chen, Lingjie
Rana, Simmi
Yang, Wendy H.
Tong, Hanghang
He, Jingrui
Information Retrieval
Understanding scientific papers requires more than answering isolated questions or summarizing content. It involves an integrated reasoning process that grounds textual and visual information, interprets experimental evidence, synthesizes information across sources, and critically evaluates scientific claims. However, existing benchmarks typically assess these abilities in isolation, making it difficult to evaluate scientific paper understanding as a unified set of interacting cognitive abilities. In this work, we introduce PaperMind, a benchmark designed to evaluate integrated and agent-oriented scientific reasoning over research papers. PaperMind is constructed from real scientific papers across seven domains, including agriculture, biology, chemistry, computer science, medicine, physics, and economics. It comprises four complementary task families that collectively operationalize distinct cognitive facets of scientific paper reasoning, including multimodal grounding, experimental interpretation, cross-source evidence reasoning, and critical assessment. By analyzing model behavior across multiple tasks, PaperMind enables a diagnostic evaluation of integrated scientific reasoning behaviors that are difficult to assess through isolated task evaluations. Extensive experiments on both opensource and closed-source multimodal LLMs reveal consistent performance gaps across tasks, highlighting persistent challenges in integrated scientific reasoning and critique. Our benchmark and dataset are available at https:// github.com/Yanjun-Zhao/PaperMind.
title PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
topic Information Retrieval
url https://arxiv.org/abs/2604.21304