Psychometric-Based Evaluation for Theorem Proving with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jianyu, Zhao, Yongwang, Zhang, Long, Hu, Jilin, Luan, Xiaokun, Xu, Zhiwei, Yang, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
by: Hu, Jilin, et al.
Published: (2025)
by: Hu, Jilin, et al.
Published: (2025)
Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4
by: Zhang, Jianyu, et al.
Published: (2026)
by: Zhang, Jianyu, et al.
Published: (2026)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025)
by: Luan, Xiaokun, et al.
Published: (2025)
Rethinking Supervision Granularity: Segment-Level Learning for LLM-Based Theorem Proving
by: Xu, Shuo, et al.
Published: (2026)
by: Xu, Shuo, et al.
Published: (2026)
Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean
by: Song, Peiyang, et al.
Published: (2024)
by: Song, Peiyang, et al.
Published: (2024)
Proving Theorems Recursively
by: Wang, Haiming, et al.
Published: (2024)
by: Wang, Haiming, et al.
Published: (2024)
Mathesis: Towards Formal Theorem Proving from Natural Languages
by: Xuejun, Yu, et al.
Published: (2025)
by: Xuejun, Yu, et al.
Published: (2025)
Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models
by: Cao, Chenrui, et al.
Published: (2025)
by: Cao, Chenrui, et al.
Published: (2025)
A Survey on Deep Learning for Theorem Proving
by: Li, Zhaoyu, et al.
Published: (2024)
by: Li, Zhaoyu, et al.
Published: (2024)
MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026)
by: Xu, Qiyuan, et al.
Published: (2026)
RLMEval: Evaluating Research-Level Neural Theorem Proving
by: Poiroux, Auguste, et al.
Published: (2025)
by: Poiroux, Auguste, et al.
Published: (2025)
PhysProver: Advancing Automatic Theorem Proving for Physics
by: Zhang, Hanning, et al.
Published: (2026)
by: Zhang, Hanning, et al.
Published: (2026)
FVEL: Interactive Formal Verification Environment with Large Language Models via Theorem Proving
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
by: Lai, Junyu, et al.
Published: (2025)
by: Lai, Junyu, et al.
Published: (2025)
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
by: Xiong, Beibei, et al.
Published: (2025)
by: Xiong, Beibei, et al.
Published: (2025)
Distilling LLM Feedback for Lean Theorem Proving
by: Narozniak, Gaetan, et al.
Published: (2026)
by: Narozniak, Gaetan, et al.
Published: (2026)
A Minimal Agent for Automated Theorem Proving
by: Requena, Borja, et al.
Published: (2026)
by: Requena, Borja, et al.
Published: (2026)
OProver: A Unified Framework for Agentic Formal Theorem Proving
by: Ma, David, et al.
Published: (2026)
by: Ma, David, et al.
Published: (2026)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
by: Chen, Luoxin, et al.
Published: (2025)
by: Chen, Luoxin, et al.
Published: (2025)
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024)
by: Hu, Jiewen, et al.
Published: (2024)
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
by: Ye, Haoran, et al.
Published: (2024)
by: Ye, Haoran, et al.
Published: (2024)
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
by: Ji, Xingguang, et al.
Published: (2025)
by: Ji, Xingguang, et al.
Published: (2025)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
by: Liu, Chengwu, et al.
Published: (2026)
by: Liu, Chengwu, et al.
Published: (2026)
FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models
by: Uluşan, Zeynel A., et al.
Published: (2026)
by: Uluşan, Zeynel A., et al.
Published: (2026)
BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving
by: Xin, Ran, et al.
Published: (2025)
by: Xin, Ran, et al.
Published: (2025)
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025)
by: Achim, Tudor, et al.
Published: (2025)
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
by: Xin, Huajian, et al.
Published: (2024)
by: Xin, Huajian, et al.
Published: (2024)
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
by: Ye, Haoran, et al.
Published: (2025)
by: Ye, Haoran, et al.
Published: (2025)
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
by: Lin, Yong, et al.
Published: (2025)
by: Lin, Yong, et al.
Published: (2025)
LeanConjecturer: Automatic Generation of Mathematical Conjectures for Theorem Proving
by: Onda, Naoto, et al.
Published: (2025)
by: Onda, Naoto, et al.
Published: (2025)
Alchemy: Amplifying Theorem-Proving Capability through Symbolic Mutation
by: Wu, Shaonan, et al.
Published: (2024)
by: Wu, Shaonan, et al.
Published: (2024)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
by: Zhao, Xueliang, et al.
Published: (2024)
by: Zhao, Xueliang, et al.
Published: (2024)
AgriGPT: a Large Language Model Ecosystem for Agriculture
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Partial Label Learning for Automated Theorem Proving
by: Zombori, Zsolt, et al.
Published: (2025)
by: Zombori, Zsolt, et al.
Published: (2025)
Similar Items
-
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
by: Hu, Jilin, et al.
Published: (2025) -
Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4
by: Zhang, Jianyu, et al.
Published: (2026) -
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025) -
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025) -
Rethinking Supervision Granularity: Segment-Level Learning for LLM-Based Theorem Proving
by: Xu, Shuo, et al.
Published: (2026)