One-Eval: An Agentic System for Automated and Traceable LLM Evaluation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shen, Chengyu, Hou, Yanheng, Pan, Minghui, He, Runming, Wong, Zhen Hao, Qiang, Meiyi, Liu, Zhou, Liang, Hao, Lai, Peichao, Sheng, Zeang, Zhang, Wentao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918381595131904
author Shen, Chengyu
Hou, Yanheng
Pan, Minghui
He, Runming
Wong, Zhen Hao
Qiang, Meiyi
Liu, Zhou
Liang, Hao
Lai, Peichao
Sheng, Zeang
Zhang, Wentao
author_facet Shen, Chengyu
Hou, Yanheng
Pan, Minghui
He, Runming
Wong, Zhen Hao
Qiang, Meiyi
Liu, Zhou
Liang, Hao
Lai, Peichao
Sheng, Zeang
Zhang, Wentao
contents Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify appropriate benchmarks, reproduce heterogeneous evaluation codebases, configure dataset schema mappings, and interpret aggregated metrics. To address these challenges, we present One-Eval, an agentic evaluation system that converts natural-language evaluation requests into executable, traceable, and customizable evaluation workflows. One-Eval integrates (i) NL2Bench for intent structuring and personalized benchmark planning, (ii) BenchResolve for benchmark resolution, automatic dataset acquisition, and schema normalization to ensure executability, and (iii) Metrics \& Reporting for task-aware metric selection and decision-oriented reporting beyond scalar scores. The system further incorporates human-in-the-loop checkpoints for review, editing, and rollback, while preserving sample evidence trails for debugging and auditability. Experiments show that One-Eval can execute end-to-end evaluations from diverse natural-language requests with minimal user effort, supporting more efficient and reproducible evaluation in industrial settings. Our framework is publicly available at https://github.com/OpenDCAI/One-Eval.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
Shen, Chengyu
Hou, Yanheng
Pan, Minghui
He, Runming
Wong, Zhen Hao
Qiang, Meiyi
Liu, Zhou
Liang, Hao
Lai, Peichao
Sheng, Zeang
Zhang, Wentao
Computation and Language
Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify appropriate benchmarks, reproduce heterogeneous evaluation codebases, configure dataset schema mappings, and interpret aggregated metrics. To address these challenges, we present One-Eval, an agentic evaluation system that converts natural-language evaluation requests into executable, traceable, and customizable evaluation workflows. One-Eval integrates (i) NL2Bench for intent structuring and personalized benchmark planning, (ii) BenchResolve for benchmark resolution, automatic dataset acquisition, and schema normalization to ensure executability, and (iii) Metrics \& Reporting for task-aware metric selection and decision-oriented reporting beyond scalar scores. The system further incorporates human-in-the-loop checkpoints for review, editing, and rollback, while preserving sample evidence trails for debugging and auditability. Experiments show that One-Eval can execute end-to-end evaluations from diverse natural-language requests with minimal user effort, supporting more efficient and reproducible evaluation in industrial settings. Our framework is publicly available at https://github.com/OpenDCAI/One-Eval.
title One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
topic Computation and Language
url https://arxiv.org/abs/2603.09821