LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ming, Peng, Qiyuan, Wei, Yinxi, Shen, Yujiong, Tan, Kexin, Wang, Yuhui, Xiang, Zhenghao, Ye, Junjie, Yin, Zhangyue, Xi, Zhiheng, Dou, Shihan, Gui, Tao, Pan, Maxm, Yang, Ruizhi, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917511014907904
author Zhang, Ming
Peng, Qiyuan
Wei, Yinxi
Shen, Yujiong
Tan, Kexin
Wang, Yuhui
Xiang, Zhenghao
Ye, Junjie
Yin, Zhangyue
Xi, Zhiheng
Dou, Shihan
Gui, Tao
Pan, Maxm
Yang, Ruizhi
Zhang, Qi
Huang, Xuanjing
author_facet Zhang, Ming
Peng, Qiyuan
Wei, Yinxi
Shen, Yujiong
Tan, Kexin
Wang, Yuhui
Xiang, Zhenghao
Ye, Junjie
Yin, Zhangyue
Xi, Zhiheng
Dou, Shihan
Gui, Tao
Pan, Maxm
Yang, Ruizhi
Zhang, Qi
Huang, Xuanjing
contents Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by templating natural-language items from sampled formulas, provide only coarse or unaudited formal annotations, and are now quickly saturated by frontier reasoning models. We present LLMEval-Logic, a Chinese logical reasoning benchmark built from realistic situational scenarios. Its pipeline forward-authors and expert-audits natural-language items together with their reference formalizations, verifies annotated answers with Z3, constructs expert rubrics for natural-to-formal grading, and hardens selected items through a closed-loop adversarial workflow. The benchmark is released in two paired subsets: a 246-item Base subset shipped with 1,400 expert-developed rubric atoms, and a 190-item Hard subset with 938 multi-step sub-questions over closed model spaces. Evaluating 14 frontier LLMs on LLMEval-Logic reveals substantial gaps in current models: the best model reaches only 37.5% Hard Item Accuracy, and even with reference symbols the highest joint Z3+Rubric formalization score among evaluated models reaches only 60.16%. Our benchmark is publicly available at https://github.com/llmeval/LLMEval-Logic.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19597
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
Zhang, Ming
Peng, Qiyuan
Wei, Yinxi
Shen, Yujiong
Tan, Kexin
Wang, Yuhui
Xiang, Zhenghao
Ye, Junjie
Yin, Zhangyue
Xi, Zhiheng
Dou, Shihan
Gui, Tao
Pan, Maxm
Yang, Ruizhi
Zhang, Qi
Huang, Xuanjing
Computation and Language
Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by templating natural-language items from sampled formulas, provide only coarse or unaudited formal annotations, and are now quickly saturated by frontier reasoning models. We present LLMEval-Logic, a Chinese logical reasoning benchmark built from realistic situational scenarios. Its pipeline forward-authors and expert-audits natural-language items together with their reference formalizations, verifies annotated answers with Z3, constructs expert rubrics for natural-to-formal grading, and hardens selected items through a closed-loop adversarial workflow. The benchmark is released in two paired subsets: a 246-item Base subset shipped with 1,400 expert-developed rubric atoms, and a 190-item Hard subset with 938 multi-step sub-questions over closed model spaces. Evaluating 14 frontier LLMs on LLMEval-Logic reveals substantial gaps in current models: the best model reaches only 37.5% Hard Item Accuracy, and even with reference symbols the highest joint Z3+Rubric formalization score among evaluated models reaches only 60.16%. Our benchmark is publicly available at https://github.com/llmeval/LLMEval-Logic.
title LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
topic Computation and Language
url https://arxiv.org/abs/2605.19597