Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Haoyu, Geng, Yihan, Tang, Shange, Lin, Yong, Lyu, Bohan, Lin, Hongzhou, Jin, Chi, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
par: Lin, Yong, et autres
Publié: (2025)
par: Lin, Yong, et autres
Publié: (2025)
Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
par: Lin, Yong, et autres
Publié: (2025)
par: Lin, Yong, et autres
Publié: (2025)
Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover
par: Chung, Jui-Hui, et autres
Publié: (2026)
par: Chung, Jui-Hui, et autres
Publié: (2026)
On Reasoning-Centric LLM-based Automated Theorem Proving
par: Sun, Yican, et autres
Publié: (2026)
par: Sun, Yican, et autres
Publié: (2026)
Benchmarking Testing in Automated Theorem Proving
par: Kim, Jongyoon, et autres
Publié: (2026)
par: Kim, Jongyoon, et autres
Publié: (2026)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
par: Chen, Luoxin, et autres
Publié: (2025)
par: Chen, Luoxin, et autres
Publié: (2025)
Is Elo Rating Reliable? A Study Under Model Misspecification
par: Tang, Shange, et autres
Publié: (2025)
par: Tang, Shange, et autres
Publié: (2025)
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
par: Xiong, Beibei, et autres
Publié: (2025)
par: Xiong, Beibei, et autres
Publié: (2025)
Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem Proving
par: Zheng, Xinyi, et autres
Publié: (2025)
par: Zheng, Xinyi, et autres
Publié: (2025)
Automated Theorem Proving for Prolog Verification
par: Mesnard, Fred, et autres
Publié: (2026)
par: Mesnard, Fred, et autres
Publié: (2026)
Canonical for Automated Theorem Proving in Lean
par: Norman, Chase, et autres
Publié: (2025)
par: Norman, Chase, et autres
Publié: (2025)
Principled Out-of-Distribution Generalization via Simplicity
par: Ge, Jiawei, et autres
Publié: (2025)
par: Ge, Jiawei, et autres
Publié: (2025)
Benign Overfitting in Out-of-Distribution Generalization of Linear Models
par: Tang, Shange, et autres
Publié: (2024)
par: Tang, Shange, et autres
Publié: (2024)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
par: Li, Zenan, et autres
Publié: (2025)
par: Li, Zenan, et autres
Publié: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
par: Cheng, Yun, et autres
Publié: (2026)
par: Cheng, Yun, et autres
Publié: (2026)
Partial Label Learning for Automated Theorem Proving
par: Zombori, Zsolt, et autres
Publié: (2025)
par: Zombori, Zsolt, et autres
Publié: (2025)
A Minimal Agent for Automated Theorem Proving
par: Requena, Borja, et autres
Publié: (2026)
par: Requena, Borja, et autres
Publié: (2026)
Aristotle: IMO-level Automated Theorem Proving
par: Achim, Tudor, et autres
Publié: (2025)
par: Achim, Tudor, et autres
Publié: (2025)
Proving Theorems Recursively
par: Wang, Haiming, et autres
Publié: (2024)
par: Wang, Haiming, et autres
Publié: (2024)
Learning to Reason with Insight for Informal Theorem Proving
par: Li, Yunhe, et autres
Publié: (2026)
par: Li, Yunhe, et autres
Publié: (2026)
Automated Discovery of Tactic Libraries for Interactive Theorem Proving
par: Xin, Yutong, et autres
Publié: (2025)
par: Xin, Yutong, et autres
Publié: (2025)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
par: Zheng, Chuanyang, et autres
Publié: (2023)
par: Zheng, Chuanyang, et autres
Publié: (2023)
Can Models Learn Skill Composition from Examples?
par: Zhao, Haoyu, et autres
Publié: (2024)
par: Zhao, Haoyu, et autres
Publié: (2024)
MSC-180: A Benchmark for Automated Formal Theorem Proving from Mathematical Subject Classification
par: Li, Sirui, et autres
Publié: (2025)
par: Li, Sirui, et autres
Publié: (2025)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
par: Wu, Yikai, et autres
Publié: (2025)
par: Wu, Yikai, et autres
Publié: (2025)
BAIT: Benchmarking (Embedding) Architectures for Interactive Theorem-Proving
par: Lamont, Sean, et autres
Publié: (2024)
par: Lamont, Sean, et autres
Publié: (2024)
Proving Olympiad Algebraic Inequalities without Human Demonstrations
par: Wei, Chenrui, et autres
Publié: (2024)
par: Wei, Chenrui, et autres
Publié: (2024)
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
par: He, Yinghui, et autres
Publié: (2025)
par: He, Yinghui, et autres
Publié: (2025)
Skill-Targeted Adaptive Training
par: He, Yinghui, et autres
Publié: (2025)
par: He, Yinghui, et autres
Publié: (2025)
Why is Your Language Model a Poor Implicit Reward Model?
par: Razin, Noam, et autres
Publié: (2025)
par: Razin, Noam, et autres
Publié: (2025)
EvolProver: Advancing Automated Theorem Proving by Evolving Formalized Problems via Symmetry and Difficulty
par: Tian, Yuchen, et autres
Publié: (2025)
par: Tian, Yuchen, et autres
Publié: (2025)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
par: Liu, Chengwu, et autres
Publié: (2026)
par: Liu, Chengwu, et autres
Publié: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
par: Petrov, Ivo, et autres
Publié: (2025)
par: Petrov, Ivo, et autres
Publié: (2025)
Reviving DSP for Advanced Theorem Proving in the Era of Reasoning Models
par: Cao, Chenrui, et autres
Publié: (2025)
par: Cao, Chenrui, et autres
Publié: (2025)
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
par: Wu, Xindi, et autres
Publié: (2024)
par: Wu, Xindi, et autres
Publié: (2024)
CompBench: Benchmarking Complex Instruction-guided Image Editing
par: Jia, Bohan, et autres
Publié: (2025)
par: Jia, Bohan, et autres
Publié: (2025)
State Canonization and Early Pruning in Width-Based Automated Theorem Proving
par: Oliveira, Mateus de Oliveira, et autres
Publié: (2026)
par: Oliveira, Mateus de Oliveira, et autres
Publié: (2026)
Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving
par: Qiu, Ruichen, et autres
Publié: (2026)
par: Qiu, Ruichen, et autres
Publié: (2026)
Self-organized criticality driven by droplet influx and random fusion
par: Lyu, Bohan, et autres
Publié: (2025)
par: Lyu, Bohan, et autres
Publié: (2025)
Analytical Calculation of Viscosity in Rouse Networks Below Gelation Transition
par: Lyu, Bohan, et autres
Publié: (2025)
par: Lyu, Bohan, et autres
Publié: (2025)
Documents similaires
-
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
par: Lin, Yong, et autres
Publié: (2025) -
Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
par: Lin, Yong, et autres
Publié: (2025) -
Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover
par: Chung, Jui-Hui, et autres
Publié: (2026) -
On Reasoning-Centric LLM-based Automated Theorem Proving
par: Sun, Yican, et autres
Publié: (2026) -
Benchmarking Testing in Automated Theorem Proving
par: Kim, Jongyoon, et autres
Publié: (2026)