CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Junqi, Lin, Xiaohan, Bayer, Jonas, Dillies, Yael, Jiang, Weijie, Liang, Xiaodan, Soletskyi, Roman, Wang, Haiming, Xie, Yunzhou, Xiong, Beibei, Yang, Zhengfeng, Zhang, Jujian, Zhi, Lihong, Li, Jia, Liu, Zhengying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Formal Proof of the Irrationality of $ζ(3)$ in Lean 4
by: Liu, Junqi, et al.
Published: (2025)
by: Liu, Junqi, et al.
Published: (2025)
Automated Formal Proofs of Combinatorial Identities via Wilf-Zeilberger Guidance and LLMs
by: Xiong, Beibei, et al.
Published: (2026)
by: Xiong, Beibei, et al.
Published: (2026)
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
by: Xiong, Beibei, et al.
Published: (2025)
by: Xiong, Beibei, et al.
Published: (2025)
LeanGeo: Formalizing Competitional Geometry problems in Lean
by: Song, Chendong, et al.
Published: (2025)
by: Song, Chendong, et al.
Published: (2025)
FVEL: Interactive Formal Verification Environment with Large Language Models via Theorem Proving
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
A Generalisation of Sperner's Theorem Using Weighted Chains
by: Dillies, Yaël, et al.
Published: (2025)
by: Dillies, Yaël, et al.
Published: (2025)
MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
by: Huang, Yinya, et al.
Published: (2024)
by: Huang, Yinya, et al.
Published: (2024)
Training Safe Neural Networks with Global SDP Bounds
by: Soletskyi, Roman, et al.
Published: (2024)
by: Soletskyi, Roman, et al.
Published: (2024)
Anchored Dyck Paths
by: Dillies, Jimmy
Published: (2026)
by: Dillies, Jimmy
Published: (2026)
CombiMOTS: Combinatorial Multi-Objective Tree Search for Dual-Target Molecule Generation
by: Southiratn, Thibaud, et al.
Published: (2026)
by: Southiratn, Thibaud, et al.
Published: (2026)
Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics
by: Liu, Junqi, et al.
Published: (2026)
by: Liu, Junqi, et al.
Published: (2026)
ATG: Benchmarking Automated Theorem Generation for Generative Language Models
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
A theoretical perspective on mode collapse in variational inference
by: Soletskyi, Roman, et al.
Published: (2024)
by: Soletskyi, Roman, et al.
Published: (2024)
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
by: Mahdavi, Hamed, et al.
Published: (2025)
by: Mahdavi, Hamed, et al.
Published: (2025)
Kimina Lean Server: A High-Performance Lean Server for Large-Scale Verification
by: Santos, Marco Dos, et al.
Published: (2025)
by: Santos, Marco Dos, et al.
Published: (2025)
Engineering of Reversibly Cross‐Linked Elastomers Toward Flexible and Recyclable Elastomer/Carbon Fiber Composites with Extraordinary Tearing Resistance
by: Xiaohan Wang, et al.
Published: (2024)
by: Xiaohan Wang, et al.
Published: (2024)
Automated Tactics for Polynomial Reasoning in Lean 4
by: Shen, Hao, et al.
Published: (2026)
by: Shen, Hao, et al.
Published: (2026)
Formalizing Gröbner Basis Theory in Lean
by: Guo, Junyu, et al.
Published: (2026)
by: Guo, Junyu, et al.
Published: (2026)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
by: Lin, Jingli, et al.
Published: (2025)
by: Lin, Jingli, et al.
Published: (2025)
CombiGCN: An effective GCN model for Recommender System
by: Nguyen, Loc Tan, et al.
Published: (2025)
by: Nguyen, Loc Tan, et al.
Published: (2025)
Formalizing multi-graded Brenner-Schröer Proj schemes and dilatations of rings in Lean4
by: Mayeux, Arnaud, et al.
Published: (2026)
by: Mayeux, Arnaud, et al.
Published: (2026)
The mechanization of science illustrated by the Lean formalization of the multi-graded Proj construction
by: Mayeux, Arnaud, et al.
Published: (2025)
by: Mayeux, Arnaud, et al.
Published: (2025)
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
by: Firsching, Moritz, et al.
Published: (2026)
by: Firsching, Moritz, et al.
Published: (2026)
Proving Theorems Recursively
by: Wang, Haiming, et al.
Published: (2024)
by: Wang, Haiming, et al.
Published: (2024)
Bayesian imaging inverse problem with scattering transform
by: Pierre, Sébastien, et al.
Published: (2026)
by: Pierre, Sébastien, et al.
Published: (2026)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
by: Jiang, Weisen, et al.
Published: (2023)
by: Jiang, Weisen, et al.
Published: (2023)
Accelerating Regression Tasks with Quantum Algorithms
by: Liu, Chenghua, et al.
Published: (2025)
by: Liu, Chenghua, et al.
Published: (2025)
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
by: Peyronnet, Antoine, et al.
Published: (2026)
by: Peyronnet, Antoine, et al.
Published: (2026)
Combi-CAM: A Novel Multi-Layer Approach for Explainable Image Geolocalization
by: Faget, David, et al.
Published: (2026)
by: Faget, David, et al.
Published: (2026)
Improvement of variables interpretability in kernel PCA
by: Briscik, Mitja, et al.
Published: (2023)
by: Briscik, Mitja, et al.
Published: (2023)
Chebyshev quotients, Demazure multiplicities, and Dyck-path models
by: Biswal, Rekha, et al.
Published: (2026)
by: Biswal, Rekha, et al.
Published: (2026)
Reciprocals of Partition Polynomials
by: Chen, Evan, et al.
Published: (2026)
by: Chen, Evan, et al.
Published: (2026)
A Formal Proof of Complexity Bounds on Diophantine Equations
by: Bayer, Jonas, et al.
Published: (2025)
by: Bayer, Jonas, et al.
Published: (2025)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
by: Yu, Longhui, et al.
Published: (2023)
by: Yu, Longhui, et al.
Published: (2023)
FormalAlign: Automated Alignment Evaluation for Autoformalization
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
Uniform Preorders and Partial Combinatory Algebras
by: Frey, Jonas
Published: (2024)
by: Frey, Jonas
Published: (2024)
Code for "Topographical Features of the Lunar Surface Unveiled through Spherical Voronoi Tessellation Analysis Method of DEM Roughness at Equal Spatial Scales "
by: Zhengfeng, Zhang
Published: (2025)
by: Zhengfeng, Zhang
Published: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
by: Tian, Runchu, et al.
Published: (2024)
by: Tian, Runchu, et al.
Published: (2024)
Similar Items
-
A Formal Proof of the Irrationality of $ζ(3)$ in Lean 4
by: Liu, Junqi, et al.
Published: (2025) -
Automated Formal Proofs of Combinatorial Identities via Wilf-Zeilberger Guidance and LLMs
by: Xiong, Beibei, et al.
Published: (2026) -
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
by: Xiong, Beibei, et al.
Published: (2025) -
LeanGeo: Formalizing Competitional Geometry problems in Lean
by: Song, Chendong, et al.
Published: (2025) -
FVEL: Interactive Formal Verification Environment with Large Language Models via Theorem Proving
by: Lin, Xiaohan, et al.
Published: (2024)