Leanabell-Prover: Posttraining Scaling in Formal Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jingyuan, Wang, Qi, Ji, Xingguang, Liu, Yahui, Yue, Yang, Zhang, Fuzheng, Zhang, Di, Zhou, Guorui, Gai, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
von: Fu, Jia, et al.
Veröffentlicht: (2025)
von: Fu, Jia, et al.
Veröffentlicht: (2025)
Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
von: Ou, Jiao, et al.
Veröffentlicht: (2024)
von: Ou, Jiao, et al.
Veröffentlicht: (2024)
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
von: Xiaomi, LLM-Core, et al.
Veröffentlicht: (2025)
von: Xiaomi, LLM-Core, et al.
Veröffentlicht: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
von: Wang, Haiming, et al.
Veröffentlicht: (2025)
von: Wang, Haiming, et al.
Veröffentlicht: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs
von: Sun, Chenxi, et al.
Veröffentlicht: (2024)
von: Sun, Chenxi, et al.
Veröffentlicht: (2024)
Chain-of-Specificity: An Iteratively Refining Method for Eliciting Knowledge from Large Language Models
von: Wei, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wei, Kaiwen, et al.
Veröffentlicht: (2024)
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
von: Lin, Lei, et al.
Veröffentlicht: (2023)
von: Lin, Lei, et al.
Veröffentlicht: (2023)
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
von: Li, Chenyi, et al.
Veröffentlicht: (2026)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
Clarifying Before Reasoning: A Coq Prover with Structural Context
von: Lu, Yanzhen, et al.
Veröffentlicht: (2025)
von: Lu, Yanzhen, et al.
Veröffentlicht: (2025)
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing
von: Sun, Chenxi, et al.
Veröffentlicht: (2025)
von: Sun, Chenxi, et al.
Veröffentlicht: (2025)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
von: Yuan, Shihao, et al.
Veröffentlicht: (2025)
von: Yuan, Shihao, et al.
Veröffentlicht: (2025)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
Prover Agent: An Agent-Based Framework for Formal Mathematical Proofs
von: Baba, Kaito, et al.
Veröffentlicht: (2025)
von: Baba, Kaito, et al.
Veröffentlicht: (2025)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
von: Lin, Yong, et al.
Veröffentlicht: (2025)
von: Lin, Yong, et al.
Veröffentlicht: (2025)
PhysProver: Advancing Automatic Theorem Proving for Physics
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
Selection and Exploitation of High-Quality Knowledge from Large Language Models for Recommendation
von: Wang, Guanchen, et al.
Veröffentlicht: (2025)
von: Wang, Guanchen, et al.
Veröffentlicht: (2025)
M2F: Automated Formalization of Mathematical Literature at Scale
von: Wang, Zichen, et al.
Veröffentlicht: (2026)
von: Wang, Zichen, et al.
Veröffentlicht: (2026)
SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin
von: Yi, Hao, et al.
Veröffentlicht: (2025)
von: Yi, Hao, et al.
Veröffentlicht: (2025)
EvolProver: Advancing Automated Theorem Proving by Evolving Formalized Problems via Symmetry and Difficulty
von: Tian, Yuchen, et al.
Veröffentlicht: (2025)
von: Tian, Yuchen, et al.
Veröffentlicht: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
StepFun-Prover Preview: Let's Think and Verify Step by Step
von: Shang, Shijie, et al.
Veröffentlicht: (2025)
von: Shang, Shijie, et al.
Veröffentlicht: (2025)
LLM-Aligned Geographic Item Tokenization for Local-Life Recommendation
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
Agentic Entropy-Balanced Policy Optimization
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
von: Li, Guchan, et al.
Veröffentlicht: (2026)
von: Li, Guchan, et al.
Veröffentlicht: (2026)
Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
von: Xin, Ran, et al.
Veröffentlicht: (2025)
von: Xin, Ran, et al.
Veröffentlicht: (2025)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
von: Li, Mukai, et al.
Veröffentlicht: (2025)
von: Li, Mukai, et al.
Veröffentlicht: (2025)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
von: Chen, Luoxin, et al.
Veröffentlicht: (2026)
von: Chen, Luoxin, et al.
Veröffentlicht: (2026)
TSO: Self-Training with Scaled Preference Optimization
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
von: Ji, Xingguang, et al.
Veröffentlicht: (2025) -
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
von: Wang, Qi, et al.
Veröffentlicht: (2025) -
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
von: Fu, Jia, et al.
Veröffentlicht: (2025) -
Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
von: Ji, Xingguang, et al.
Veröffentlicht: (2025) -
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)