FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xiao-Wen, Zhang, Zihao, Cao, Jianuo, Zhou, Zhi, Li, Zenan, Guo, Lan-Zhe, Yao, Yuan, Chen, Taolue, Li, Yu-Feng, Ma, Xiaoxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Disprove: Formal Counterexample Generation with Large Language Models
by: Li, Zenan, et al.
Published: (2026)
by: Li, Zenan, et al.
Published: (2026)
Softened Symbol Grounding for Neuro-symbolic Systems
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
Conformal Correction for Efficiency May be at Odds with Entropy
by: Xu, Senrong, et al.
Published: (2025)
by: Xu, Senrong, et al.
Published: (2025)
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
by: Yousefzadeh, Roozbeh, et al.
Published: (2025)
by: Yousefzadeh, Roozbeh, et al.
Published: (2025)
Neuro-Symbolic Proof Generation for Scaling Systems Software Verification
by: He, Baoding, et al.
Published: (2026)
by: He, Baoding, et al.
Published: (2026)
Learning with Logical Constraints but without Shortcut Satisfaction
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Task Abstention for Large Language Models in Code Generation
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
Fair Conformal Classification via Learning Representation-Based Groups
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Neuro-symbolic Learning Yielding Logical Constraints
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Formalizing Gröbner Basis Theory in Lean
by: Guo, Junyu, et al.
Published: (2026)
by: Guo, Junyu, et al.
Published: (2026)
Neuro-Symbolic Data Generation for Math Reasoning
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
Synthesizing Inductive Invariants for Distributed Protocols via IC3 and Large Language Models
by: Cao, Weining, et al.
Published: (2026)
by: Cao, Weining, et al.
Published: (2026)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
by: Yu, Zhouliang, et al.
Published: (2025)
by: Yu, Zhouliang, et al.
Published: (2025)
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
by: Ren, Z. Z., et al.
Published: (2025)
by: Ren, Z. Z., et al.
Published: (2025)
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
by: Zhang, Gehao, et al.
Published: (2026)
by: Zhang, Gehao, et al.
Published: (2026)
An Algorithm for Diagonalizing Matrices of Formal Power Series
by: Dai, Zihao, et al.
Published: (2026)
by: Dai, Zihao, et al.
Published: (2026)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving
by: Qiu, Ruichen, et al.
Published: (2026)
by: Qiu, Ruichen, et al.
Published: (2026)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
by: Dong, Yihong, et al.
Published: (2026)
by: Dong, Yihong, et al.
Published: (2026)
Formalizing Wu-Ritt Method in Lean 4
by: Xiao, Yuxuan, et al.
Published: (2026)
by: Xiao, Yuxuan, et al.
Published: (2026)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
by: Huang, Yunpeng, et al.
Published: (2023)
by: Huang, Yunpeng, et al.
Published: (2023)
Formal description of ML models for unambiguous implementation
by: Gauffriau, Adrien, et al.
Published: (2023)
by: Gauffriau, Adrien, et al.
Published: (2023)
FormalASR: End-to-End Spoken Chinese to Formal Text
by: Ning, Wanyi, et al.
Published: (2026)
by: Ning, Wanyi, et al.
Published: (2026)
Formal P-Category Theory and Normalization by Evaluation in Rocq
by: Berry, David G., et al.
Published: (2025)
by: Berry, David G., et al.
Published: (2025)
DRAMPyML: A Formal Description of DRAM Protocols with Timed Petri Nets
by: Christ, Derek, et al.
Published: (2026)
by: Christ, Derek, et al.
Published: (2026)
Formal Theory at ICHEP 2024
by: Schafer-Nameki, Sakura
Published: (2024)
by: Schafer-Nameki, Sakura
Published: (2024)
The Formal Theory of Monads, Univalently
by: van der Weide, Niels
Published: (2022)
by: van der Weide, Niels
Published: (2022)
Formalization of Amicable Numbers Theory
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
Cypher is Turing-Complete: A Formal Proof via 2-Counter Machine Simulation
by: Halftermeyer, Pierre
Published: (2026)
by: Halftermeyer, Pierre
Published: (2026)
FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models
by: Uluşan, Zeynel A., et al.
Published: (2026)
by: Uluşan, Zeynel A., et al.
Published: (2026)
LeanCat: A Benchmark Suite for Formal Category Theory in Lean (Part I: 1-Categories)
by: Xu, Rongge, et al.
Published: (2025)
by: Xu, Rongge, et al.
Published: (2025)
FormalGeo: An Extensible Formalized Framework for Olympiad Geometric Problem Solving
by: Zhang, Xiaokai, et al.
Published: (2023)
by: Zhang, Xiaokai, et al.
Published: (2023)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
by: Li, Zenan, et al.
Published: (2025)
by: Li, Zenan, et al.
Published: (2025)
SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
by: Cheng, Qian, et al.
Published: (2025)
by: Cheng, Qian, et al.
Published: (2025)
A Formal Theory of Survey Experiment Generalizability: Attention and Salience
by: Fu, Jiawei, et al.
Published: (2024)
by: Fu, Jiawei, et al.
Published: (2024)
Formalizing Poisson-Boltzmann Theory for Field-Tunable Nanofluidic Devices
by: Zhao, Zhongyuan, et al.
Published: (2026)
by: Zhao, Zhongyuan, et al.
Published: (2026)
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
by: Xu, Xu, et al.
Published: (2025)
by: Xu, Xu, et al.
Published: (2025)
Similar Items
-
Learning to Disprove: Formal Counterexample Generation with Large Language Models
by: Li, Zenan, et al.
Published: (2026) -
Softened Symbol Grounding for Neuro-symbolic Systems
by: Li, Zenan, et al.
Published: (2024) -
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025) -
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025) -
Conformal Correction for Efficiency May be at Odds with Entropy
by: Xu, Senrong, et al.
Published: (2025)