RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yelin, Zhang, Fanjin, Sun, Suping, Pang, Yunhe, Wang, Yuanchun, Song, Jian, Li, Xiaoyan, Hou, Lei, Zhao, Shu, Tang, Jie, Li, Juanzi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915969601896448
author Chen, Yelin
Zhang, Fanjin
Sun, Suping
Pang, Yunhe
Wang, Yuanchun
Song, Jian
Li, Xiaoyan
Hou, Lei
Zhao, Shu
Tang, Jie
Li, Juanzi
author_facet Chen, Yelin
Zhang, Fanjin
Sun, Suping
Pang, Yunhe
Wang, Yuanchun
Song, Jian
Li, Xiaoyan
Hou, Lei
Zhao, Shu
Tang, Jie
Li, Juanzi
contents Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we introduce RPC-Bench, a large-scale question-answering benchmark built from review-rebuttal exchanges of high-quality computer science papers, containing 15K human-verified QA pairs. We design a fine-grained taxonomy aligned with the scientific research flow to assess models' ability to understand and answer why, what, and how questions in scholarly contexts. We also define an elaborate LLM-human interaction annotation framework to support large-scale labeling and quality control. Following the LLM-as-a-Judge paradigm, we develop a scalable framework that evaluates models on correctness-completeness and conciseness, with high agreement to human judgment. Experiments reveal that even the strongest models (GPT-5) achieve only 68.2% correctness-completeness, dropping to 37.46% after conciseness adjustment, highlighting substantial gaps in precise academic paper understanding. Our code and data are available at https://rpc-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14289
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
Chen, Yelin
Zhang, Fanjin
Sun, Suping
Pang, Yunhe
Wang, Yuanchun
Song, Jian
Li, Xiaoyan
Hou, Lei
Zhao, Shu
Tang, Jie
Li, Juanzi
Computation and Language
Artificial Intelligence
Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we introduce RPC-Bench, a large-scale question-answering benchmark built from review-rebuttal exchanges of high-quality computer science papers, containing 15K human-verified QA pairs. We design a fine-grained taxonomy aligned with the scientific research flow to assess models' ability to understand and answer why, what, and how questions in scholarly contexts. We also define an elaborate LLM-human interaction annotation framework to support large-scale labeling and quality control. Following the LLM-as-a-Judge paradigm, we develop a scalable framework that evaluates models on correctness-completeness and conciseness, with high agreement to human judgment. Experiments reveal that even the strongest models (GPT-5) achieve only 68.2% correctness-completeness, dropping to 37.46% after conciseness adjustment, highlighting substantial gaps in precise academic paper understanding. Our code and data are available at https://rpc-bench.github.io/.
title RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.14289