DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xin, Huajian, Ren, Z. Z., Song, Junxiao, Shao, Zhihong, Zhao, Wanjia, Wang, Haocheng, Liu, Bo, Zhang, Liyue, Lu, Xuan, Du, Qiushi, Gao, Wenjun, Zhu, Qihao, Yang, Dejian, Gou, Zhibin, Wu, Z. F., Luo, Fuli, Ruan, Chong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
von: Guo, Daya, et al.
Veröffentlicht: (2024)
von: Guo, Daya, et al.
Veröffentlicht: (2024)
DeepSeek-V3 Technical Report
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition
von: Tang, Haocheng, et al.
Veröffentlicht: (2026)
von: Tang, Haocheng, et al.
Veröffentlicht: (2026)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
von: DeepSeek-AI, et al.
Veröffentlicht: (2025)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
Multi-Prover Interactive Proof Systems with Leakage
von: Asadi, Vahid R., et al.
Veröffentlicht: (2026)
von: Asadi, Vahid R., et al.
Veröffentlicht: (2026)
A Comparison of DeepSeek and Other LLMs
von: Gao, Tianchen, et al.
Veröffentlicht: (2025)
von: Gao, Tianchen, et al.
Veröffentlicht: (2025)
DeepSeek-OCR: Contexts Optical Compression
von: Wei, Haoran, et al.
Veröffentlicht: (2025)
von: Wei, Haoran, et al.
Veröffentlicht: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
von: Shao, Zhihong, et al.
Veröffentlicht: (2024)
von: Shao, Zhihong, et al.
Veröffentlicht: (2024)
DeepSeek reshaping healthcare in China's tertiary hospitals
von: Chen, Jishizhan, et al.
Veröffentlicht: (2025)
von: Chen, Jishizhan, et al.
Veröffentlicht: (2025)
Safety Evaluation of DeepSeek Models in Chinese Contexts
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
DeepSeek-OCR 2: Visual Causal Flow
von: Wei, Haoran, et al.
Veröffentlicht: (2026)
von: Wei, Haoran, et al.
Veröffentlicht: (2026)
Proof Recommendation System for the HOL4 Theorem Prover
von: Dekhil, Nour, et al.
Veröffentlicht: (2024)
von: Dekhil, Nour, et al.
Veröffentlicht: (2024)
Proof Strategy Extraction from LLMs for Enhancing Symbolic Provers
von: Fang, Jian, et al.
Veröffentlicht: (2025)
von: Fang, Jian, et al.
Veröffentlicht: (2025)
MerLean-Prover: A Recursive Looping Harness for Lean 4 Theorem Proving
von: Li, Jinzheng, et al.
Veröffentlicht: (2026)
von: Li, Jinzheng, et al.
Veröffentlicht: (2026)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
von: Xiao, Dongfu, et al.
Veröffentlicht: (2025)
von: Xiao, Dongfu, et al.
Veröffentlicht: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
An evaluation of DeepSeek Models in Biomedical Natural Language Processing
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2025)
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
von: Ahuja, Riyaz, et al.
Veröffentlicht: (2026)
von: Ahuja, Riyaz, et al.
Veröffentlicht: (2026)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
Proof Assistants for Teaching: a Survey
von: Minh, Frédéric Tran, et al.
Veröffentlicht: (2025)
von: Minh, Frédéric Tran, et al.
Veröffentlicht: (2025)
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
von: Zhang, Chong, et al.
Veröffentlicht: (2025)
Institutional Trust and the Domestic AI Advantage: Evidence from DeepSeek and ChatGPT Users in China
von: Huang, Jiashen, et al.
Veröffentlicht: (2026)
von: Huang, Jiashen, et al.
Veröffentlicht: (2026)
Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
von: Spelda, Petr, et al.
Veröffentlicht: (2025)
von: Spelda, Petr, et al.
Veröffentlicht: (2025)
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
von: Li, Xuying, et al.
Veröffentlicht: (2025)
von: Li, Xuying, et al.
Veröffentlicht: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Innovating China's Intangible Cultural Heritage with DeepSeek + MidJourney: The Case of Yangliuqing theme Woodblock Prints
von: Yang, RuiKun, et al.
Veröffentlicht: (2025)
von: Yang, RuiKun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
von: Ren, Z. Z., et al.
Veröffentlicht: (2025) -
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
von: Xin, Huajian, et al.
Veröffentlicht: (2024) -
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
von: Shao, Zhihong, et al.
Veröffentlicht: (2025) -
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
von: Guo, Daya, et al.
Veröffentlicht: (2024) -
DeepSeek-V3 Technical Report
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)