Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Chuxue, Li, Mengze, Dai, Juntao, Yang, Jinluan, Zhao, Zijian, Zhang, Shengyu, Shi, Weijie, Liu, Chengzhong, Han, Sirui, Guo, Yike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912442978664448
author Cao, Chuxue
Li, Mengze
Dai, Juntao
Yang, Jinluan
Zhao, Zijian
Zhang, Shengyu
Shi, Weijie
Liu, Chengzhong
Han, Sirui
Guo, Yike
author_facet Cao, Chuxue
Li, Mengze
Dai, Juntao
Yang, Jinluan
Zhao, Zijian
Zhang, Shengyu
Shi, Weijie
Liu, Chengzhong
Han, Sirui
Guo, Yike
contents Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas. However, their effectiveness in complex mathematical reasoning involving multi-step FOL deductions is still under-researched. While LLMs perform competitively on established mathematical reasoning benchmarks, they struggle with multi-step FOL tasks, as demonstrated by Deepseek-Prover-V2-7B's low accuracy (4.2%) on our proposed theorem proving dataset. This issue arises from the limited exploration of diverse proof strategies and the potential for early reasoning mistakes to undermine entire proofs. To address these issues, we propose DREAM, a self-adaptive solution that enhances the Diversity and REAsonability of LLMs' generation strategies. DREAM incorporates an Axiom-Driven Strategy Diversification mechanism to promote varied strategic outcomes and a Sub-Proposition Error Feedback to help LLMs reflect on and correct their proofs. Our contributions include pioneering advancements in LLMs' mathematical reasoning through FOL theorem proving, introducing a novel inference stage solution that improves performance by 0.6% to 6.4%, and providing a curated dataset of 447 mathematical theorems in Lean 4 format for evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
Cao, Chuxue
Li, Mengze
Dai, Juntao
Yang, Jinluan
Zhao, Zijian
Zhang, Shengyu
Shi, Weijie
Liu, Chengzhong
Han, Sirui
Guo, Yike
Artificial Intelligence
Computation and Language
Logic in Computer Science
Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas. However, their effectiveness in complex mathematical reasoning involving multi-step FOL deductions is still under-researched. While LLMs perform competitively on established mathematical reasoning benchmarks, they struggle with multi-step FOL tasks, as demonstrated by Deepseek-Prover-V2-7B's low accuracy (4.2%) on our proposed theorem proving dataset. This issue arises from the limited exploration of diverse proof strategies and the potential for early reasoning mistakes to undermine entire proofs. To address these issues, we propose DREAM, a self-adaptive solution that enhances the Diversity and REAsonability of LLMs' generation strategies. DREAM incorporates an Axiom-Driven Strategy Diversification mechanism to promote varied strategic outcomes and a Sub-Proposition Error Feedback to help LLMs reflect on and correct their proofs. Our contributions include pioneering advancements in LLMs' mathematical reasoning through FOL theorem proving, introducing a novel inference stage solution that improves performance by 0.6% to 6.4%, and providing a curated dataset of 447 mathematical theorems in Lean 4 format for evaluation.
title Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
topic Artificial Intelligence
Computation and Language
Logic in Computer Science
url https://arxiv.org/abs/2506.17104