An Evaluation Benchmark for Autoformalization in Lean4
Fuente:
arXiv
Guardado en:
| Autores principales: | Gulati, Aryan, Ladsaria, Devanshu, Mishra, Shubhra, Sidhu, Jasdeep, Miranda, Brando |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
por: Chan, Willy, et al.
Publicado: (2025)
por: Chan, Willy, et al.
Publicado: (2025)
Quantifying the Importance of Data Alignment in Downstream Model Performance
por: Chawla, Krrish, et al.
Publicado: (2025)
por: Chawla, Krrish, et al.
Publicado: (2025)
Reliable Evaluation and Benchmarks for Statement Autoformalization
por: Poiroux, Auguste, et al.
Publicado: (2024)
por: Poiroux, Auguste, et al.
Publicado: (2024)
FormalAlign: Automated Alignment Evaluation for Autoformalization
por: Lu, Jianqiao, et al.
Publicado: (2024)
por: Lu, Jianqiao, et al.
Publicado: (2024)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
por: Agarwal, Anmol, et al.
Publicado: (2026)
por: Agarwal, Anmol, et al.
Publicado: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
por: Zhou, Jin Peng, et al.
Publicado: (2024)
por: Zhou, Jin Peng, et al.
Publicado: (2024)
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
por: Liu, Xiaoyang, et al.
Publicado: (2025)
por: Liu, Xiaoyang, et al.
Publicado: (2025)
Process-Driven Autoformalization in Lean 4
por: Lu, Jianqiao, et al.
Publicado: (2024)
por: Lu, Jianqiao, et al.
Publicado: (2024)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
por: Wu, Yutong, et al.
Publicado: (2025)
por: Wu, Yutong, et al.
Publicado: (2025)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
por: Steinmetz, Cody, et al.
Publicado: (2025)
por: Steinmetz, Cody, et al.
Publicado: (2025)
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
por: Heakl, Ahmed, et al.
Publicado: (2025)
por: Heakl, Ahmed, et al.
Publicado: (2025)
Illuminate: A novel approach for depression detection with explainable analysis and proactive therapy using prompt engineering
por: Agrawal, Aryan
Publicado: (2024)
por: Agrawal, Aryan
Publicado: (2024)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
por: Jadon, Aryan, et al.
Publicado: (2025)
por: Jadon, Aryan, et al.
Publicado: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
por: Obbad, Elyas, et al.
Publicado: (2024)
por: Obbad, Elyas, et al.
Publicado: (2024)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Evaluating LLMs for Hardware Design and Test
por: Blocklove, Jason, et al.
Publicado: (2024)
por: Blocklove, Jason, et al.
Publicado: (2024)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
por: Dong, Honghua, et al.
Publicado: (2024)
por: Dong, Honghua, et al.
Publicado: (2024)
AIOS Compiler: LLM as Interpreter for Natural Language Programming and Flow Programming of AI Agents
por: Xu, Shuyuan, et al.
Publicado: (2024)
por: Xu, Shuyuan, et al.
Publicado: (2024)
Code Simulation Challenges for Large Language Models
por: La Malfa, Emanuele, et al.
Publicado: (2024)
por: La Malfa, Emanuele, et al.
Publicado: (2024)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
por: Ravi, Nikil, et al.
Publicado: (2026)
por: Ravi, Nikil, et al.
Publicado: (2026)
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Bridging the Knowledge Void: Inference-time Acquisition of Unfamiliar Programming Languages for Coding Tasks
por: Shen, Chen, et al.
Publicado: (2026)
por: Shen, Chen, et al.
Publicado: (2026)
LILO: Learning Interpretable Libraries by Compressing and Documenting Code
por: Grand, Gabriel, et al.
Publicado: (2023)
por: Grand, Gabriel, et al.
Publicado: (2023)
Emergent Representations of Program Semantics in Language Models Trained on Programs
por: Jin, Charles, et al.
Publicado: (2023)
por: Jin, Charles, et al.
Publicado: (2023)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
por: Tsoukalas, George, et al.
Publicado: (2024)
por: Tsoukalas, George, et al.
Publicado: (2024)
Can't Remember Details in Long Documents? You Need Some R&R
por: Agrawal, Devanshu, et al.
Publicado: (2024)
por: Agrawal, Devanshu, et al.
Publicado: (2024)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
por: Zhang, Ke, et al.
Publicado: (2026)
por: Zhang, Ke, et al.
Publicado: (2026)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
por: Schaeffer, Rylan, et al.
Publicado: (2024)
por: Schaeffer, Rylan, et al.
Publicado: (2024)
Prakriti200: A Questionnaire-Based Dataset of 200 Ayurvedic Prakriti Assessments
por: Singh, Aryan Kumar, et al.
Publicado: (2025)
por: Singh, Aryan Kumar, et al.
Publicado: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
por: Xin, Yutong, et al.
Publicado: (2026)
por: Xin, Yutong, et al.
Publicado: (2026)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
por: Sajith, Aryan, et al.
Publicado: (2024)
por: Sajith, Aryan, et al.
Publicado: (2024)
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
por: Lalwani, Abhinav, et al.
Publicado: (2024)
por: Lalwani, Abhinav, et al.
Publicado: (2024)
Neural Task Synthesis for Visual Programming
por: Pădurean, Victor-Alexandru, et al.
Publicado: (2023)
por: Pădurean, Victor-Alexandru, et al.
Publicado: (2023)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
por: Zhang, Yike, et al.
Publicado: (2025)
por: Zhang, Yike, et al.
Publicado: (2025)
CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models
por: Kalyan, Pavan, et al.
Publicado: (2025)
por: Kalyan, Pavan, et al.
Publicado: (2025)
Large Language Models aren't all that you need
por: Holla, Kiran Voderhobli, et al.
Publicado: (2024)
por: Holla, Kiran Voderhobli, et al.
Publicado: (2024)
BEExAI: Benchmark to Evaluate Explainable AI
por: Sithakoul, Samuel, et al.
Publicado: (2024)
por: Sithakoul, Samuel, et al.
Publicado: (2024)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
por: Labrak, Yanis, et al.
Publicado: (2024)
por: Labrak, Yanis, et al.
Publicado: (2024)
Ejemplares similares
-
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
por: Chan, Willy, et al.
Publicado: (2025) -
Quantifying the Importance of Data Alignment in Downstream Model Performance
por: Chawla, Krrish, et al.
Publicado: (2025) -
Reliable Evaluation and Benchmarks for Statement Autoformalization
por: Poiroux, Auguste, et al.
Publicado: (2024) -
FormalAlign: Automated Alignment Evaluation for Autoformalization
por: Lu, Jianqiao, et al.
Publicado: (2024) -
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
por: Agarwal, Anmol, et al.
Publicado: (2026)