The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dekoninck, Jasper, Petrov, Ivo, Minchev, Kristian, Balunovic, Mislav, Vechev, Martin, Marinov, Miroslav, Drencheva, Maria, Konova, Lyuba, Shumanov, Milen, Tsvetkov, Kaloyan, Drenchev, Nikolay, Todorov, Lazar, Nikolova, Kalina, Georgiev, Nikolay, Kalinkova, Vanesa, Ismoldayev, Margulan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
CuTS: Customizable Tabular Synthetic Data Generation
von: Vero, Mark, et al.
Veröffentlicht: (2023)
von: Vero, Mark, et al.
Veröffentlicht: (2023)
Large Language Models are Advanced Anonymizers
von: Staab, Robin, et al.
Veröffentlicht: (2024)
von: Staab, Robin, et al.
Veröffentlicht: (2024)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
von: Staab, Robin, et al.
Veröffentlicht: (2023)
von: Staab, Robin, et al.
Veröffentlicht: (2023)
ToolFuzz -- Automated Agent Tool Testing
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
GRAIN: Exact Graph Reconstruction from Gradients
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
von: Drencheva, Maria, et al.
Veröffentlicht: (2025)
BV photometry of the ultracompact binary star GP Com
von: Zamanov, Radoslav, et al.
Veröffentlicht: (2026)
von: Zamanov, Radoslav, et al.
Veröffentlicht: (2026)
Proof of the Complete Presence of a Modulo 4 Bias for the Semiprimes
von: Gyulev, Nikola, et al.
Veröffentlicht: (2024)
von: Gyulev, Nikola, et al.
Veröffentlicht: (2024)
AlphaIntegrator: Transformer Action Search for Symbolic Integration Proofs
von: Ünsal, Mert, et al.
Veröffentlicht: (2024)
von: Ünsal, Mert, et al.
Veröffentlicht: (2024)
A Unified Approach to Routing and Cascading for LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
von: Hiss, Hanno, et al.
Veröffentlicht: (2026)
von: Hiss, Hanno, et al.
Veröffentlicht: (2026)
ConStat: Performance-Based Contamination Detection in Large Language Models
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
Specific features of the $π$-electron spectrum of narrow achiral $(2m,m)$ nanoribbons
von: Malysheva, Lyuba
Veröffentlicht: (2025)
von: Malysheva, Lyuba
Veröffentlicht: (2025)
Adaptive Generation of Bias-Eliciting Questions for LLMs
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
Passive uplift of montane biotas: recent advances
von: Michael Heads, et al.
Veröffentlicht: (2025)
von: Michael Heads, et al.
Veröffentlicht: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
von: Debenedetti, Edoardo, et al.
Veröffentlicht: (2024)
Controlled Text Generation via Language Model Arithmetic
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning
von: Ho, Matthew, et al.
Veröffentlicht: (2024)
von: Ho, Matthew, et al.
Veröffentlicht: (2024)
Working apart together: How outcome control supports operations when working from home
von: Henri C. Dekker, et al.
Veröffentlicht: (2025)
von: Henri C. Dekker, et al.
Veröffentlicht: (2025)
Proofs that Modify Proofs
von: Towsner, Henry
Veröffentlicht: (2024)
von: Towsner, Henry
Veröffentlicht: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
All Proof of Work But No Proof of Play
von: Tirmazi, Hayder
Veröffentlicht: (2025)
von: Tirmazi, Hayder
Veröffentlicht: (2025)
Proofs that Modify Proofs, 1/2
von: Towsner, Henry
Veröffentlicht: (2025)
von: Towsner, Henry
Veröffentlicht: (2025)
Symmetric Proofs in the Ideal Proof System
von: Dawar, Anuj, et al.
Veröffentlicht: (2025)
von: Dawar, Anuj, et al.
Veröffentlicht: (2025)
Beyond the lone hero: How interpersonal feedback seeking helps entrepreneurs to engage with their social environment
von: Andreana Drencheva, et al.
Veröffentlicht: (2024)
von: Andreana Drencheva, et al.
Veröffentlicht: (2024)
TECHNOLOGICAL POSSIBILITIES FOR THE BENEFICIATION OF OXIDE ORES FROM A PORPHYRY COPPER DEPOSIT
von: Yankova, Teodora, et al.
Veröffentlicht: (2025)
von: Yankova, Teodora, et al.
Veröffentlicht: (2025)
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis
von: Petrov, Nikolay B, et al.
Veröffentlicht: (2024)
von: Petrov, Nikolay B, et al.
Veröffentlicht: (2024)
Unravelling Abstract Cyclic Proofs into Proofs by Induction
von: Grotenhuis, Lide, et al.
Veröffentlicht: (2026)
von: Grotenhuis, Lide, et al.
Veröffentlicht: (2026)
ProofCloud: A Proof Retrieval Engine for Verified Proofs in Higher Order Logic
von: Wang, Shuai
Veröffentlicht: (2024)
von: Wang, Shuai
Veröffentlicht: (2024)
How to do safe and shorter transabdominal preperitoneal ( TAPP) inguinal hernia repair
von: Kiril Georgiev Kirov, et al.
Veröffentlicht: (2024)
von: Kiril Georgiev Kirov, et al.
Veröffentlicht: (2024)
New mineralogical and geochemical data on a potential critical raw materials occurrence at Polski Gradets (Southеastern Bulgaria)
von: Stavrev, Milen, et al.
Veröffentlicht: (2024)
von: Stavrev, Milen, et al.
Veröffentlicht: (2024)
Punctually Standard and Nonstandard Models of Natural Numbers
von: Bazhenov, Nikolay, et al.
Veröffentlicht: (2026)
von: Bazhenov, Nikolay, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025) -
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
von: Balunović, Mislav, et al.
Veröffentlicht: (2025) -
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025) -
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026) -
CuTS: Customizable Tabular Synthetic Data Generation
von: Vero, Mark, et al.
Veröffentlicht: (2023)