LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917511014907904 |
|---|---|
| author | Zhang, Ming Peng, Qiyuan Wei, Yinxi Shen, Yujiong Tan, Kexin Wang, Yuhui Xiang, Zhenghao Ye, Junjie Yin, Zhangyue Xi, Zhiheng Dou, Shihan Gui, Tao Pan, Maxm Yang, Ruizhi Zhang, Qi Huang, Xuanjing |
| author_facet | Zhang, Ming Peng, Qiyuan Wei, Yinxi Shen, Yujiong Tan, Kexin Wang, Yuhui Xiang, Zhenghao Ye, Junjie Yin, Zhangyue Xi, Zhiheng Dou, Shihan Gui, Tao Pan, Maxm Yang, Ruizhi Zhang, Qi Huang, Xuanjing |
| contents | Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by templating natural-language items from sampled formulas, provide only coarse or unaudited formal annotations, and are now quickly saturated by frontier reasoning models. We present LLMEval-Logic, a Chinese logical reasoning benchmark built from realistic situational scenarios. Its pipeline forward-authors and expert-audits natural-language items together with their reference formalizations, verifies annotated answers with Z3, constructs expert rubrics for natural-to-formal grading, and hardens selected items through a closed-loop adversarial workflow. The benchmark is released in two paired subsets: a 246-item Base subset shipped with 1,400 expert-developed rubric atoms, and a 190-item Hard subset with 938 multi-step sub-questions over closed model spaces. Evaluating 14 frontier LLMs on LLMEval-Logic reveals substantial gaps in current models: the best model reaches only 37.5% Hard Item Accuracy, and even with reference symbols the highest joint Z3+Rubric formalization score among evaluated models reaches only 60.16%. Our benchmark is publicly available at https://github.com/llmeval/LLMEval-Logic. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_19597 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Zhang, Ming Peng, Qiyuan Wei, Yinxi Shen, Yujiong Tan, Kexin Wang, Yuhui Xiang, Zhenghao Ye, Junjie Yin, Zhangyue Xi, Zhiheng Dou, Shihan Gui, Tao Pan, Maxm Yang, Ruizhi Zhang, Qi Huang, Xuanjing Computation and Language Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow strictly from stated premises. Many existing logical-reasoning benchmarks are generated by templating natural-language items from sampled formulas, provide only coarse or unaudited formal annotations, and are now quickly saturated by frontier reasoning models. We present LLMEval-Logic, a Chinese logical reasoning benchmark built from realistic situational scenarios. Its pipeline forward-authors and expert-audits natural-language items together with their reference formalizations, verifies annotated answers with Z3, constructs expert rubrics for natural-to-formal grading, and hardens selected items through a closed-loop adversarial workflow. The benchmark is released in two paired subsets: a 246-item Base subset shipped with 1,400 expert-developed rubric atoms, and a 190-item Hard subset with 938 multi-step sub-questions over closed model spaces. Evaluating 14 frontier LLMs on LLMEval-Logic reveals substantial gaps in current models: the best model reaches only 37.5% Hard Item Accuracy, and even with reference symbols the highest joint Z3+Rubric formalization score among evaluated models reaches only 60.16%. Our benchmark is publicly available at https://github.com/llmeval/LLMEval-Logic. |
| title | LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2605.19597 |