SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Yilun, Zhang, Kaiyan, Hu, Tiansheng, Wu, Sihong, Bras, Ronan Le, McGrady, Charles, Anderson, Taira, Bragg, Jonathan, Chang, Joseph Chee, Dodge, Jesse, Latzke, Matt, Liu, Yixin, Tang, Xiangru, Wang, Zihang, Zhao, Chen, Hajishirzi, Hannaneh, Downey, Doug, Cohan, Arman
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917217867661312
author Zhao, Yilun
Zhang, Kaiyan
Hu, Tiansheng
Wu, Sihong
Bras, Ronan Le
McGrady, Charles
Anderson, Taira
Bragg, Jonathan
Chang, Joseph Chee
Dodge, Jesse
Latzke, Matt
Liu, Yixin
Tang, Xiangru
Wang, Zihang
Zhao, Chen
Hajishirzi, Hannaneh
Downey, Doug
Cohan, Arman
author_facet Zhao, Yilun
Zhang, Kaiyan
Hu, Tiansheng
Wu, Sihong
Bras, Ronan Le
McGrady, Charles
Anderson, Taira
Bragg, Jonathan
Chang, Joseph Chee
Dodge, Jesse
Latzke, Matt
Liu, Yixin
Tang, Xiangru
Wang, Zihang
Zhao, Chen
Hajishirzi, Hannaneh
Downey, Doug
Cohan, Arman
contents We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses. The platform currently supports 47 foundation models and has collected over 20,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality. We discuss the results and insights based on the model ranking leaderboard. To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on collected preference data. It measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark's challenges and emphasize the need for more reliable automated evaluation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Zhao, Yilun
Zhang, Kaiyan
Hu, Tiansheng
Wu, Sihong
Bras, Ronan Le
McGrady, Charles
Anderson, Taira
Bragg, Jonathan
Chang, Joseph Chee
Dodge, Jesse
Latzke, Matt
Liu, Yixin
Tang, Xiangru
Wang, Zihang
Zhao, Chen
Hajishirzi, Hannaneh
Downey, Doug
Cohan, Arman
Computation and Language
Artificial Intelligence
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses. The platform currently supports 47 foundation models and has collected over 20,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality. We discuss the results and insights based on the model ranking leaderboard. To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on collected preference data. It measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark's challenges and emphasize the need for more reliable automated evaluation methods.
title SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.01001