Mathematics with large language models as provers and verifiers
Fuente:
arXiv
Salvato in:
| Autori principali: | Duc, Hieu Le, Liberti, Leo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
di: Noël, Valentin
Pubblicazione: (2026)
di: Noël, Valentin
Pubblicazione: (2026)
Bolzano: Case Studies in LLM-Assisted Mathematical Research
di: Balko, Martin, et al.
Pubblicazione: (2026)
di: Balko, Martin, et al.
Pubblicazione: (2026)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
di: Shen, Ziju, et al.
Pubblicazione: (2025)
di: Shen, Ziju, et al.
Pubblicazione: (2025)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
di: Tsoukalas, George, et al.
Pubblicazione: (2024)
di: Tsoukalas, George, et al.
Pubblicazione: (2024)
Large language models as oracles for instantiating ontologies with domain-specific knowledge
di: Ciatto, Giovanni, et al.
Pubblicazione: (2024)
di: Ciatto, Giovanni, et al.
Pubblicazione: (2024)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
di: Chen, Michael K., et al.
Pubblicazione: (2025)
di: Chen, Michael K., et al.
Pubblicazione: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
Transformers Can Learn Connectivity in Some Graphs but Not Others
di: Roy, Amit, et al.
Pubblicazione: (2025)
di: Roy, Amit, et al.
Pubblicazione: (2025)
Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
di: Bao, Qiming, et al.
Pubblicazione: (2025)
di: Bao, Qiming, et al.
Pubblicazione: (2025)
Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving
di: Rajaee, Sara, et al.
Pubblicazione: (2025)
di: Rajaee, Sara, et al.
Pubblicazione: (2025)
Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
di: Arakelyan, Erik
Pubblicazione: (2025)
di: Arakelyan, Erik
Pubblicazione: (2025)
Combining Textual and Structural Information for Premise Selection in Lean
di: Petrovčič, Job, et al.
Pubblicazione: (2025)
di: Petrovčič, Job, et al.
Pubblicazione: (2025)
Are Language Models Efficient Reasoners? A Perspective from Logic Programming
di: Opedal, Andreas, et al.
Pubblicazione: (2025)
di: Opedal, Andreas, et al.
Pubblicazione: (2025)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
di: Su, DiJia, et al.
Pubblicazione: (2025)
di: Su, DiJia, et al.
Pubblicazione: (2025)
A Neurosymbolic Approach to Natural Language Formalization and Verification
di: Bayless, Sam, et al.
Pubblicazione: (2025)
di: Bayless, Sam, et al.
Pubblicazione: (2025)
The Geometry of Reasoning: Flowing Logics in Representation Space
di: Zhou, Yufa, et al.
Pubblicazione: (2025)
di: Zhou, Yufa, et al.
Pubblicazione: (2025)
Self-Supervised Transformers as Iterative Solution Improvers for Constraint Satisfaction
di: Xu, Yudong W., et al.
Pubblicazione: (2025)
di: Xu, Yudong W., et al.
Pubblicazione: (2025)
Hierarchical Attention Generates Better Proofs
di: Chen, Jianlong, et al.
Pubblicazione: (2025)
di: Chen, Jianlong, et al.
Pubblicazione: (2025)
Recursive Decomposition of Logical Thoughts: Framework for Superior Reasoning and Knowledge Propagation in Large Language Models
di: Qasim, Kaleem Ullah, et al.
Pubblicazione: (2025)
di: Qasim, Kaleem Ullah, et al.
Pubblicazione: (2025)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
di: Zhao, Xueliang, et al.
Pubblicazione: (2024)
di: Zhao, Xueliang, et al.
Pubblicazione: (2024)
RLSF: Fine-tuning LLMs via Symbolic Feedback
di: Jha, Piyush, et al.
Pubblicazione: (2024)
di: Jha, Piyush, et al.
Pubblicazione: (2024)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
di: Lalwani, Abhinav, et al.
Pubblicazione: (2024)
di: Lalwani, Abhinav, et al.
Pubblicazione: (2024)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
di: Bao, Qiming, et al.
Pubblicazione: (2022)
di: Bao, Qiming, et al.
Pubblicazione: (2022)
Consistent Joint Decision-Making with Heterogeneous Learning Models
di: Faghihi, Hossein Rajaby, et al.
Pubblicazione: (2024)
di: Faghihi, Hossein Rajaby, et al.
Pubblicazione: (2024)
Controlling Logical Collapse in LLMs via Algebraic Ontology Projection over F2
di: Miyashita, Hisashi, et al.
Pubblicazione: (2026)
di: Miyashita, Hisashi, et al.
Pubblicazione: (2026)
Guiding Word Equation Solving using Graph Neural Networks (Extended Technical Report)
di: Abdulla, Parosh Aziz, et al.
Pubblicazione: (2024)
di: Abdulla, Parosh Aziz, et al.
Pubblicazione: (2024)
FLARE: Faithful Logic-Aided Reasoning and Exploration
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
di: Ahuja, Riyaz, et al.
Pubblicazione: (2026)
di: Ahuja, Riyaz, et al.
Pubblicazione: (2026)
Quantifying artificial intelligence through algorithmic generalization
di: Ito, Takuya, et al.
Pubblicazione: (2024)
di: Ito, Takuya, et al.
Pubblicazione: (2024)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
di: Xin, Huajian, et al.
Pubblicazione: (2024)
di: Xin, Huajian, et al.
Pubblicazione: (2024)
DeepOnto: A Python Package for Ontology Engineering with Deep Learning
di: He, Yuan, et al.
Pubblicazione: (2023)
di: He, Yuan, et al.
Pubblicazione: (2023)
Harnessing the Power of Semi-Structured Knowledge and LLMs with Triplet-Based Prefiltering for Question Answering
di: Boer, Derian, et al.
Pubblicazione: (2024)
di: Boer, Derian, et al.
Pubblicazione: (2024)
Herald: A Natural Language Annotated Lean 4 Dataset
di: Gao, Guoxiong, et al.
Pubblicazione: (2024)
di: Gao, Guoxiong, et al.
Pubblicazione: (2024)
Formal Mathematical Reasoning: A New Frontier in AI
di: Yang, Kaiyu, et al.
Pubblicazione: (2024)
di: Yang, Kaiyu, et al.
Pubblicazione: (2024)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
di: Liu, Xinyu, et al.
Pubblicazione: (2026)
From Symbolic Tasks to Code Generation: Diversification Yields Better Task Performers
di: Zhang, Dylan, et al.
Pubblicazione: (2024)
di: Zhang, Dylan, et al.
Pubblicazione: (2024)
NLP Verification: Towards a General Methodology for Certifying Robustness
di: Casadio, Marco, et al.
Pubblicazione: (2024)
di: Casadio, Marco, et al.
Pubblicazione: (2024)
MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Llemma: An Open Language Model For Mathematics
di: Azerbayev, Zhangir, et al.
Pubblicazione: (2023)
di: Azerbayev, Zhangir, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
di: Noël, Valentin
Pubblicazione: (2026) -
Bolzano: Case Studies in LLM-Assisted Mathematical Research
di: Balko, Martin, et al.
Pubblicazione: (2026) -
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
di: Shen, Ziju, et al.
Pubblicazione: (2025) -
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
di: Tsoukalas, George, et al.
Pubblicazione: (2024) -
Large language models as oracles for instantiating ontologies with domain-specific knowledge
di: Ciatto, Giovanni, et al.
Pubblicazione: (2024)