PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Akyürek, Afra Feyza, Gosai, Advait, Zhang, Chen Bo Calvin, Gupta, Vipul, Jeong, Jaehwan, Gunjal, Anisha, Rabbani, Tahseen, Mazzone, Maria, Randolph, David, Meymand, Mohammad Mahmoudi, Chattha, Gurshaan, Rodriguez, Paula, Mares, Diego, Singh, Pavit, Liu, Michael, Chawla, Subodh, Cline, Pete, Ogaz, Lucy, Hernandez, Ernesto, Wang, Zihao, Bhatter, Pavi, Ayestaran, Marcos, Liu, Bing, He, Yunzhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
by: Tang, Zilu, et al.
Published: (2025)
by: Tang, Zilu, et al.
Published: (2025)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024)
by: Akyürek, Afra Feyza, et al.
Published: (2024)
Agentic Rubrics as Contextual Verifiers for SWE Agents
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
Reward Hacking in Rubric-Based Reinforcement Learning
by: Mahmoud, Anas, et al.
Published: (2026)
by: Mahmoud, Anas, et al.
Published: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025)
by: Gunjal, Anisha, et al.
Published: (2025)
Detecting and Preventing Hallucinations in Large Vision Language Models
by: Gunjal, Anisha, et al.
Published: (2023)
by: Gunjal, Anisha, et al.
Published: (2023)
Automated ensemble method for pediatric brain tumor segmentation
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
Balancing Label Imbalance in Federated Environments Using Only Mixup and Artificially-Labeled Noise
by: Sang, Kyle, et al.
Published: (2024)
by: Sang, Kyle, et al.
Published: (2024)
VARIOS
by: Mónica Ogáz
Published: (2011)
by: Mónica Ogáz
Published: (2011)
Scaling Laws for Neural Material Models
by: Trikha, Akshay, et al.
Published: (2025)
by: Trikha, Akshay, et al.
Published: (2025)
PRBench: A Standardized Probabilistic Robustness Benchmark
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Bermúdez Soto, Jorge (2014): Fundamentos de Derecho Ambiental (Valparaíso, Ediciones Universitarias de Valparaíso, Pontificia Universidad Católica de Valparaíso), 552 pp.
by: Teresita González Ogaz
Published: (2016)
by: Teresita González Ogaz
Published: (2016)
PRBench: End-to-end Paper Reproduction in Physics Research
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Federation over Text: Insight Sharing for Multi-Agent Reasoning
by: Yao, Dixi, et al.
Published: (2026)
by: Yao, Dixi, et al.
Published: (2026)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
A Model of Knowledge-sharing for the 21st Century Organizations
by: Sabino Ayestarán
Published: (2022)
by: Sabino Ayestarán
Published: (2022)
Ciencia, responsabilidad cosmopolita y sostenibilidad en un mundo global
by: Ignacio Ayestarán
Published: (2011)
by: Ignacio Ayestarán
Published: (2011)
Epistemología de la innovación social y de la destrucción creativa
by: Ignacio Ayestarán
Published: (2011)
by: Ignacio Ayestarán
Published: (2011)
Pensamiento abismal y ecología de saberes ante la ecuación de la modernidad. En homenaje a la obra de Boaventura de Sousa Santos
by: Ignacio Ayestarán
Published: (2011)
by: Ignacio Ayestarán
Published: (2011)
Resistencia de vigas esbeltas de acero inoxidable bajo cargas concentradas mediante elementos finitos
by: Asdrubal Ayestarán
Published: (2017)
by: Asdrubal Ayestarán
Published: (2017)
La dialéctica como contribución para el desarrollo del pensamiento
by: Leonardo Gabriel Ogaz Arce
Published: (2012)
by: Leonardo Gabriel Ogaz Arce
Published: (2012)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
by: Srinivasa, Rakshith S, et al.
Published: (2025)
by: Srinivasa, Rakshith S, et al.
Published: (2025)
Changes in colour and mechanical properties of wood polypropylene composites on natural weathering
by: Jayashri Gunjal
Published: (2020)
by: Jayashri Gunjal
Published: (2020)
Sketch-GNN: Scalable Graph Neural Networks with Sublinear Training Complexity
by: Ding, Mucong, et al.
Published: (2024)
by: Ding, Mucong, et al.
Published: (2024)
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
by: Nath, Vaskar, et al.
Published: (2025)
by: Nath, Vaskar, et al.
Published: (2025)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
by: Kumar, Deepak
Published: (2026)
by: Kumar, Deepak
Published: (2026)
Residue Constraints in the Rank-Three Lifting Problem for Projective-Plane Incidence Matrices
by: Kim, Jaehwan
Published: (2026)
by: Kim, Jaehwan
Published: (2026)
Three-Sign Cancellation Hypernumber Systems and Associator Curvature
by: Kim, Jaehwan
Published: (2026)
by: Kim, Jaehwan
Published: (2026)
Real‐Time Secondary Animation with Spring Decomposed Skinning
by: B. Akyürek, et al.
Published: (2025)
by: B. Akyürek, et al.
Published: (2025)
Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
conv_einsum: A Framework for Representation and Fast Evaluation of Multilinear Operations in Convolutional Tensorial Neural Networks
by: Rabbani, Tahseen, et al.
Published: (2024)
by: Rabbani, Tahseen, et al.
Published: (2024)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
by: Bossy, Thierry, et al.
Published: (2025)
by: Bossy, Thierry, et al.
Published: (2025)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
by: Li, Zongze, et al.
Published: (2026)
by: Li, Zongze, et al.
Published: (2026)
Asynchronous Heavy-Tailed Optimization
by: Sun, Junfei, et al.
Published: (2026)
by: Sun, Junfei, et al.
Published: (2026)
FORMULATION AND IN VITRO EVALUATION OF BUCCAL TABLETS OF LOSARTAN POTASSIUM
by: Sameena Tahseen
Published: (2025)
by: Sameena Tahseen
Published: (2025)
CUSTOMER PERCEPTION OF E-BANKING SERVICE IN INDIAN BANKS : A COMPARATIVE STUDY OF TRADITIONAL BANKING AND E-BANKING
by: Nazhat, Tahseen
Published: (2026)
by: Nazhat, Tahseen
Published: (2026)
Similar Items
-
Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
by: Tang, Zilu, et al.
Published: (2025) -
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
by: Akyürek, Afra Feyza, et al.
Published: (2024) -
Agentic Rubrics as Contextual Verifiers for SWE Agents
by: Raghavendra, Mohit, et al.
Published: (2026) -
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025) -
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)