PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akyürek, Afra Feyza, Gosai, Advait, Zhang, Chen Bo Calvin, Gupta, Vipul, Jeong, Jaehwan, Gunjal, Anisha, Rabbani, Tahseen, Mazzone, Maria, Randolph, David, Meymand, Mohammad Mahmoudi, Chattha, Gurshaan, Rodriguez, Paula, Mares, Diego, Singh, Pavit, Liu, Michael, Chawla, Subodh, Cline, Pete, Ogaz, Lucy, Hernandez, Ernesto, Wang, Zihao, Bhatter, Pavi, Ayestaran, Marcos, Liu, Bing, He, Yunzhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917080367890432
author Akyürek, Afra Feyza
Gosai, Advait
Zhang, Chen Bo Calvin
Gupta, Vipul
Jeong, Jaehwan
Gunjal, Anisha
Rabbani, Tahseen
Mazzone, Maria
Randolph, David
Meymand, Mohammad Mahmoudi
Chattha, Gurshaan
Rodriguez, Paula
Mares, Diego
Singh, Pavit
Liu, Michael
Chawla, Subodh
Cline, Pete
Ogaz, Lucy
Hernandez, Ernesto
Wang, Zihao
Bhatter, Pavi
Ayestaran, Marcos
Liu, Bing
He, Yunzhong
author_facet Akyürek, Afra Feyza
Gosai, Advait
Zhang, Chen Bo Calvin
Gupta, Vipul
Jeong, Jaehwan
Gunjal, Anisha
Rabbani, Tahseen
Mazzone, Maria
Randolph, David
Meymand, Mohammad Mahmoudi
Chattha, Gurshaan
Rodriguez, Paula
Mares, Diego
Singh, Pavit
Liu, Michael
Chawla, Subodh
Cline, Pete
Ogaz, Lucy
Hernandez, Ernesto
Wang, Zihao
Bhatter, Pavi
Ayestaran, Marcos
Liu, Bing
He, Yunzhong
contents Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically consequential tasks in high-stakes domains like Legal and Finance, where practical returns are paramount. To address this, we introduce Professional Reasoning Bench (PRBench), a realistic, open-ended, and difficult benchmark of real-world problems in Finance and Law. We open-source its 1,100 expert-authored tasks and 19,356 expert-curated criteria, making it, to our knowledge, the largest public, rubric-based benchmark for both legal and finance domains. We recruit 182 qualified professionals, holding JDs, CFAs, or 6+ years of experience, who contributed tasks inspired by their actual workflows. This process yields significant diversity, with tasks spanning 114 countries and 47 US jurisdictions. Our expert-curated rubrics are validated through a rigorous quality pipeline, including independent expert validation. Subsequent evaluation of 20 leading models reveals substantial room for improvement, with top scores of only 0.39 (Finance) and 0.37 (Legal) on our Hard subsets. We further catalog associated economic impacts of the prompts and analyze performance using human-annotated rubric categories. Our analysis shows that models with similar overall scores can diverge significantly on specific capabilities. Common failure modes include inaccurate judgments, a lack of process transparency and incomplete reasoning, highlighting critical gaps in their reliability for professional adoption.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
Akyürek, Afra Feyza
Gosai, Advait
Zhang, Chen Bo Calvin
Gupta, Vipul
Jeong, Jaehwan
Gunjal, Anisha
Rabbani, Tahseen
Mazzone, Maria
Randolph, David
Meymand, Mohammad Mahmoudi
Chattha, Gurshaan
Rodriguez, Paula
Mares, Diego
Singh, Pavit
Liu, Michael
Chawla, Subodh
Cline, Pete
Ogaz, Lucy
Hernandez, Ernesto
Wang, Zihao
Bhatter, Pavi
Ayestaran, Marcos
Liu, Bing
He, Yunzhong
Computation and Language
Computers and Society
Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically consequential tasks in high-stakes domains like Legal and Finance, where practical returns are paramount. To address this, we introduce Professional Reasoning Bench (PRBench), a realistic, open-ended, and difficult benchmark of real-world problems in Finance and Law. We open-source its 1,100 expert-authored tasks and 19,356 expert-curated criteria, making it, to our knowledge, the largest public, rubric-based benchmark for both legal and finance domains. We recruit 182 qualified professionals, holding JDs, CFAs, or 6+ years of experience, who contributed tasks inspired by their actual workflows. This process yields significant diversity, with tasks spanning 114 countries and 47 US jurisdictions. Our expert-curated rubrics are validated through a rigorous quality pipeline, including independent expert validation. Subsequent evaluation of 20 leading models reveals substantial room for improvement, with top scores of only 0.39 (Finance) and 0.37 (Legal) on our Hard subsets. We further catalog associated economic impacts of the prompts and analyze performance using human-annotated rubric categories. Our analysis shows that models with similar overall scores can diverge significantly on specific capabilities. Common failure modes include inaccurate judgments, a lack of process transparency and incomplete reasoning, highlighting critical gaps in their reliability for professional adoption.
title PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2511.11562