Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Maosong, Chen, Kai, Duan, Haodong, Fang, Yixiao, Fei, Zhiwei, Gao, Tong, Jiaye, Ge, Li, Mo, Liu, Hongwei, Liu, Junnan, Liu, Yuan, Lyu, Chengqi, Lyu, Han, Ma, Ningsheng, Ma, Zerun, Sun, Yu, Wu, Zhiyong, Xiao, Linchen, Xu, Jun, Ye, Haochen, Yu, Zhaohui, Yuan, Yike, Zhang, Songyang, Zhao, Yufeng, Zhou, Fengzhe, Zhou, Peiheng, Zhu, Dongsheng, Zhu, Lin, Zhuo, Jingming
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2605.19276
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
Sommario:
  • In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.