Saved in:
Bibliographic Details
Main Authors: Ji, Yuliang, Shen, Fuchen, Wu, Jian, Xie, Qiujie, Zhang, Yue
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.20973
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915814050889728
author Ji, Yuliang
Shen, Fuchen
Wu, Jian
Xie, Qiujie
Zhang, Yue
author_facet Ji, Yuliang
Shen, Fuchen
Wu, Jian
Xie, Qiujie
Zhang, Yue
contents To comprehensively evaluate the mathematical reasoning capabilities of Large Language Models (LLMs), researchers have introduced abundant mathematical reasoning datasets. However, most existing datasets primarily focus on linear reasoning, neglecting other parts such as proof by contradiction and proof by cases, which are crucial for investigating LLMs' reasoning abilities. To address this limitation, we first introduce a novel first-order logic (FOL) dataset named PC-FOL, annotated by professional mathematicians, focusing on case-based reasoning problems. All instances in this dataset are equipped with a manually written natural language proof, clearly distinguishing it from conventional linear reasoning datasets. Our experimental results over leading LLMs demonstrate a substantial performance gap between linear reasoning and case-based reasoning problems. To further investigate this phenomenon, we provide a theoretical analysis grounded in graphical model, which provides an explanation for the observed disparity between the two types of reasoning problems. We hope this work can reveal the core challenges in the field of automated natural language mathematical proof generation, paving the way for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20973
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Linear Reasoning vs. Proof by Cases: Obstacles for Large Language Models in FOL Problem Solving
Ji, Yuliang
Shen, Fuchen
Wu, Jian
Xie, Qiujie
Zhang, Yue
Computation and Language
To comprehensively evaluate the mathematical reasoning capabilities of Large Language Models (LLMs), researchers have introduced abundant mathematical reasoning datasets. However, most existing datasets primarily focus on linear reasoning, neglecting other parts such as proof by contradiction and proof by cases, which are crucial for investigating LLMs' reasoning abilities. To address this limitation, we first introduce a novel first-order logic (FOL) dataset named PC-FOL, annotated by professional mathematicians, focusing on case-based reasoning problems. All instances in this dataset are equipped with a manually written natural language proof, clearly distinguishing it from conventional linear reasoning datasets. Our experimental results over leading LLMs demonstrate a substantial performance gap between linear reasoning and case-based reasoning problems. To further investigate this phenomenon, we provide a theoretical analysis grounded in graphical model, which provides an explanation for the observed disparity between the two types of reasoning problems. We hope this work can reveal the core challenges in the field of automated natural language mathematical proof generation, paving the way for future research.
title Linear Reasoning vs. Proof by Cases: Obstacles for Large Language Models in FOL Problem Solving
topic Computation and Language
url https://arxiv.org/abs/2602.20973