Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.19276 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918528750190592 |
|---|---|
| author | Cao, Maosong Chen, Kai Duan, Haodong Fang, Yixiao Fei, Zhiwei Gao, Tong Jiaye, Ge Li, Mo Liu, Hongwei Liu, Junnan Liu, Yuan Lyu, Chengqi Lyu, Han Ma, Ningsheng Ma, Zerun Sun, Yu Wu, Zhiyong Xiao, Linchen Xu, Jun Ye, Haochen Yu, Zhaohui Yuan, Yike Zhang, Songyang Zhao, Yufeng Zhou, Fengzhe Zhou, Peiheng Zhu, Dongsheng Zhu, Lin Zhuo, Jingming |
| author_facet | Cao, Maosong Chen, Kai Duan, Haodong Fang, Yixiao Fei, Zhiwei Gao, Tong Jiaye, Ge Li, Mo Liu, Hongwei Liu, Junnan Liu, Yuan Lyu, Chengqi Lyu, Han Ma, Ningsheng Ma, Zerun Sun, Yu Wu, Zhiyong Xiao, Linchen Xu, Jun Ye, Haochen Yu, Zhaohui Yuan, Yike Zhang, Songyang Zhao, Yufeng Zhou, Fengzhe Zhou, Peiheng Zhu, Dongsheng Zhu, Lin Zhuo, Jingming |
| contents | In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_19276 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | OpenCompass: A Universal Evaluation Platform for Large Language Models Cao, Maosong Chen, Kai Duan, Haodong Fang, Yixiao Fei, Zhiwei Gao, Tong Jiaye, Ge Li, Mo Liu, Hongwei Liu, Junnan Liu, Yuan Lyu, Chengqi Lyu, Han Ma, Ningsheng Ma, Zerun Sun, Yu Wu, Zhiyong Xiao, Linchen Xu, Jun Ye, Haochen Yu, Zhaohui Yuan, Yike Zhang, Songyang Zhao, Yufeng Zhou, Fengzhe Zhou, Peiheng Zhu, Dongsheng Zhu, Lin Zhuo, Jingming Computation and Language Machine Learning In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization. |
| title | OpenCompass: A Universal Evaluation Platform for Large Language Models |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2605.19276 |