HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhai, Weiqi, Wang, Zhihai, Wang, Jinghang, Yang, Boyu, Li, Xiaogang, Xu, Xander, Wang, Bohan, Wang, Peng, Wu, Xingzhe, Li, Anfeng, Feng, Qiyuan, Zhou, Yuhao, Han, Shoulin, Luo, Wenjie, Li, Yiyuan, Wang, Yaxuan, Luo, Ruixian, Lin, Guojie, Xiao, Peiyao, Xu, Chengliang, Wang, Ben, Wang, Zeyu, Chen, Zichao, Ye, Jianan, Hu, Yijie, Chen, Jialong, Shen, Zongwen, Xu, Yuliang, Yang, An, Yu, Bowen, Liu, Dayiheng, Lin, Junyang, Wei, Hu, Shen, Que, Zhao, Bing
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917299536003072
author Zhai, Weiqi
Wang, Zhihai
Wang, Jinghang
Yang, Boyu
Li, Xiaogang
Xu, Xander
Wang, Bohan
Wang, Peng
Wu, Xingzhe
Li, Anfeng
Feng, Qiyuan
Zhou, Yuhao
Han, Shoulin
Luo, Wenjie
Li, Yiyuan
Wang, Yaxuan
Luo, Ruixian
Lin, Guojie
Xiao, Peiyao
Xu, Chengliang
Wang, Ben
Wang, Zeyu
Chen, Zichao
Ye, Jianan
Hu, Yijie
Chen, Jialong
Shen, Zongwen
Xu, Yuliang
Yang, An
Yu, Bowen
Liu, Dayiheng
Lin, Junyang
Wei, Hu
Shen, Que
Zhao, Bing
author_facet Zhai, Weiqi
Wang, Zhihai
Wang, Jinghang
Yang, Boyu
Li, Xiaogang
Xu, Xander
Wang, Bohan
Wang, Peng
Wu, Xingzhe
Li, Anfeng
Feng, Qiyuan
Zhou, Yuhao
Han, Shoulin
Luo, Wenjie
Li, Yiyuan
Wang, Yaxuan
Luo, Ruixian
Lin, Guojie
Xiao, Peiyao
Xu, Chengliang
Wang, Ben
Wang, Zeyu
Chen, Zichao
Ye, Jianan
Hu, Yijie
Chen, Jialong
Shen, Zongwen
Xu, Yuliang
Yang, An
Yu, Bowen
Liu, Dayiheng
Lin, Junyang
Wei, Hu
Shen, Que
Zhao, Bing
contents Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified
format Preprint
id arxiv_https___arxiv_org_abs_2602_13964
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
Zhai, Weiqi
Wang, Zhihai
Wang, Jinghang
Yang, Boyu
Li, Xiaogang
Xu, Xander
Wang, Bohan
Wang, Peng
Wu, Xingzhe
Li, Anfeng
Feng, Qiyuan
Zhou, Yuhao
Han, Shoulin
Luo, Wenjie
Li, Yiyuan
Wang, Yaxuan
Luo, Ruixian
Lin, Guojie
Xiao, Peiyao
Xu, Chengliang
Wang, Ben
Wang, Zeyu
Chen, Zichao
Ye, Jianan
Hu, Yijie
Chen, Jialong
Shen, Zongwen
Xu, Yuliang
Yang, An
Yu, Bowen
Liu, Dayiheng
Lin, Junyang
Wei, Hu
Shen, Que
Zhao, Bing
Computation and Language
Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified
title HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
topic Computation and Language
url https://arxiv.org/abs/2602.13964