Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cao, Maosong, Chen, Kai, Duan, Haodong, Fang, Yixiao, Fei, Zhiwei, Gao, Tong, Jiaye, Ge, Li, Mo, Liu, Hongwei, Liu, Junnan, Liu, Yuan, Lyu, Chengqi, Lyu, Han, Ma, Ningsheng, Ma, Zerun, Sun, Yu, Wu, Zhiyong, Xiao, Linchen, Xu, Jun, Ye, Haochen, Yu, Zhaohui, Yuan, Yike, Zhang, Songyang, Zhao, Yufeng, Zhou, Fengzhe, Zhou, Peiheng, Zhu, Dongsheng, Zhu, Lin, Zhuo, Jingming
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2605.19276
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918528750190592
author Cao, Maosong
Chen, Kai
Duan, Haodong
Fang, Yixiao
Fei, Zhiwei
Gao, Tong
Jiaye, Ge
Li, Mo
Liu, Hongwei
Liu, Junnan
Liu, Yuan
Lyu, Chengqi
Lyu, Han
Ma, Ningsheng
Ma, Zerun
Sun, Yu
Wu, Zhiyong
Xiao, Linchen
Xu, Jun
Ye, Haochen
Yu, Zhaohui
Yuan, Yike
Zhang, Songyang
Zhao, Yufeng
Zhou, Fengzhe
Zhou, Peiheng
Zhu, Dongsheng
Zhu, Lin
Zhuo, Jingming
author_facet Cao, Maosong
Chen, Kai
Duan, Haodong
Fang, Yixiao
Fei, Zhiwei
Gao, Tong
Jiaye, Ge
Li, Mo
Liu, Hongwei
Liu, Junnan
Liu, Yuan
Lyu, Chengqi
Lyu, Han
Ma, Ningsheng
Ma, Zerun
Sun, Yu
Wu, Zhiyong
Xiao, Linchen
Xu, Jun
Ye, Haochen
Yu, Zhaohui
Yuan, Yike
Zhang, Songyang
Zhao, Yufeng
Zhou, Fengzhe
Zhou, Peiheng
Zhu, Dongsheng
Zhu, Lin
Zhuo, Jingming
contents In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19276
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OpenCompass: A Universal Evaluation Platform for Large Language Models
Cao, Maosong
Chen, Kai
Duan, Haodong
Fang, Yixiao
Fei, Zhiwei
Gao, Tong
Jiaye, Ge
Li, Mo
Liu, Hongwei
Liu, Junnan
Liu, Yuan
Lyu, Chengqi
Lyu, Han
Ma, Ningsheng
Ma, Zerun
Sun, Yu
Wu, Zhiyong
Xiao, Linchen
Xu, Jun
Ye, Haochen
Yu, Zhaohui
Yuan, Yike
Zhang, Songyang
Zhao, Yufeng
Zhou, Fengzhe
Zhou, Peiheng
Zhu, Dongsheng
Zhu, Lin
Zhuo, Jingming
Computation and Language
Machine Learning
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
title OpenCompass: A Universal Evaluation Platform for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.19276