CuDIP: Enhancing Theorem Proving in LLMs via Curriculum Learning-based Direct Preference Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Shuming, Zuo, Ruobing, He, Gaolei, Wang, Jianlin, Xu, Chenyang, Yang, Zhengfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From LLM-Generated Conjectures to Lean Formalizations: Automated Polynomial Inequality Proving via Sum-of-Squares Certificates
por: Zuo, Ruobing, et al.
Publicado: (2026)
por: Zuo, Ruobing, et al.
Publicado: (2026)
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
por: Xiong, Beibei, et al.
Publicado: (2025)
por: Xiong, Beibei, et al.
Publicado: (2025)
A Context-Enhanced Framework for Sequential Graph Reasoning
por: Shi, Shuo, et al.
Publicado: (2024)
por: Shi, Shuo, et al.
Publicado: (2024)
Steering LLMs for Formal Theorem Proving
por: Kirtania, Shashank, et al.
Publicado: (2025)
por: Kirtania, Shashank, et al.
Publicado: (2025)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
por: Cao, Chuxue, et al.
Publicado: (2025)
por: Cao, Chuxue, et al.
Publicado: (2025)
On Reasoning-Centric LLM-based Automated Theorem Proving
por: Sun, Yican, et al.
Publicado: (2026)
por: Sun, Yican, et al.
Publicado: (2026)
Learning to Reason with Insight for Informal Theorem Proving
por: Li, Yunhe, et al.
Publicado: (2026)
por: Li, Yunhe, et al.
Publicado: (2026)
Proving Theorems Recursively
por: Wang, Haiming, et al.
Publicado: (2024)
por: Wang, Haiming, et al.
Publicado: (2024)
Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
por: Wen, Xin-Cheng, et al.
Publicado: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
por: Fu, Yuhan, et al.
Publicado: (2024)
por: Fu, Yuhan, et al.
Publicado: (2024)
A Survey on Deep Learning for Theorem Proving
por: Li, Zhaoyu, et al.
Publicado: (2024)
por: Li, Zhaoyu, et al.
Publicado: (2024)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
por: Zhao, Xueliang, et al.
Publicado: (2024)
por: Zhao, Xueliang, et al.
Publicado: (2024)
Curriculum Direct Preference Optimization for Diffusion and Consistency Models
por: Croitoru, Florinel-Alin, et al.
Publicado: (2024)
por: Croitoru, Florinel-Alin, et al.
Publicado: (2024)
Formal Theorem Proving by Rewarding LLMs to Decompose Proofs Hierarchically
por: Dong, Kefan, et al.
Publicado: (2024)
por: Dong, Kefan, et al.
Publicado: (2024)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
por: Petrov, Ivo, et al.
Publicado: (2025)
por: Petrov, Ivo, et al.
Publicado: (2025)
Sponsored Search Auction Design Beyond Single Utility Maximization
por: Xu, Changfeng, et al.
Publicado: (2024)
por: Xu, Changfeng, et al.
Publicado: (2024)
Direct Preference Optimization for LLM-Enhanced Recommendation Systems
por: Sun, Chao, et al.
Publicado: (2024)
por: Sun, Chao, et al.
Publicado: (2024)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
por: Lai, Junyu, et al.
Publicado: (2025)
por: Lai, Junyu, et al.
Publicado: (2025)
De Novo Molecular Design Enabled by Direct Preference Optimization and Curriculum Learning
por: Hou, Junyu
Publicado: (2025)
por: Hou, Junyu
Publicado: (2025)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
por: Zhao, Mengjie, et al.
Publicado: (2026)
por: Zhao, Mengjie, et al.
Publicado: (2026)
Twitch: Learning Abstractions for Equational Theorem Proving
por: Axelrod, Guy, et al.
Publicado: (2026)
por: Axelrod, Guy, et al.
Publicado: (2026)
Partial Label Learning for Automated Theorem Proving
por: Zombori, Zsolt, et al.
Publicado: (2025)
por: Zombori, Zsolt, et al.
Publicado: (2025)
An In-Context Learning Agent for Formal Theorem-Proving
por: Thakur, Amitayush, et al.
Publicado: (2023)
por: Thakur, Amitayush, et al.
Publicado: (2023)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
por: Zhang, Ziyin, et al.
Publicado: (2025)
por: Zhang, Ziyin, et al.
Publicado: (2025)
Aligning CodeLLMs with Direct Preference Optimization
por: Miao, Yibo, et al.
Publicado: (2024)
por: Miao, Yibo, et al.
Publicado: (2024)
Automated Formal Proofs of Combinatorial Identities via Wilf-Zeilberger Guidance and LLMs
por: Xiong, Beibei, et al.
Publicado: (2026)
por: Xiong, Beibei, et al.
Publicado: (2026)
Open-Book Neural Algorithmic Reasoning
por: Li, Hefei, et al.
Publicado: (2024)
por: Li, Hefei, et al.
Publicado: (2024)
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
por: Ji, Xingguang, et al.
Publicado: (2025)
por: Ji, Xingguang, et al.
Publicado: (2025)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
por: Zhang, Xuemiao, et al.
Publicado: (2025)
por: Zhang, Xuemiao, et al.
Publicado: (2025)
2D-Curri-DPO: Two-Dimensional Curriculum Learning for Direct Preference Optimization
por: Li, Mengyang, et al.
Publicado: (2025)
por: Li, Mengyang, et al.
Publicado: (2025)
Rethinking Supervision Granularity: Segment-Level Learning for LLM-Based Theorem Proving
por: Xu, Shuo, et al.
Publicado: (2026)
por: Xu, Shuo, et al.
Publicado: (2026)
Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
por: Yang, Sen, et al.
Publicado: (2024)
por: Yang, Sen, et al.
Publicado: (2024)
LeanAgent: Lifelong Learning for Formal Theorem Proving
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
por: Kumarappan, Adarsh, et al.
Publicado: (2024)
Psychometric-Based Evaluation for Theorem Proving with Large Language Models
por: Zhang, Jianyu, et al.
Publicado: (2025)
por: Zhang, Jianyu, et al.
Publicado: (2025)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
por: Yang, Xiao, et al.
Publicado: (2024)
por: Yang, Xiao, et al.
Publicado: (2024)
OProver: A Unified Framework for Agentic Formal Theorem Proving
por: Ma, David, et al.
Publicado: (2026)
por: Ma, David, et al.
Publicado: (2026)
Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation
por: Croitoru, Florinel-Alin, et al.
Publicado: (2026)
por: Croitoru, Florinel-Alin, et al.
Publicado: (2026)
Benchmarking Testing in Automated Theorem Proving
por: Kim, Jongyoon, et al.
Publicado: (2026)
por: Kim, Jongyoon, et al.
Publicado: (2026)
Automated Theorem Proving for Prolog Verification
por: Mesnard, Fred, et al.
Publicado: (2026)
por: Mesnard, Fred, et al.
Publicado: (2026)
Canonical for Automated Theorem Proving in Lean
por: Norman, Chase, et al.
Publicado: (2025)
por: Norman, Chase, et al.
Publicado: (2025)
Ejemplares similares
-
From LLM-Generated Conjectures to Lean Formalizations: Automated Polynomial Inequality Proving via Sum-of-Squares Certificates
por: Zuo, Ruobing, et al.
Publicado: (2026) -
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation
por: Xiong, Beibei, et al.
Publicado: (2025) -
A Context-Enhanced Framework for Sequential Graph Reasoning
por: Shi, Shuo, et al.
Publicado: (2024) -
Steering LLMs for Formal Theorem Proving
por: Kirtania, Shashank, et al.
Publicado: (2025) -
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
por: Cao, Chuxue, et al.
Publicado: (2025)