Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lyu, Chengqi, Gao, Songyang, Gu, Yuzhe, Zhang, Wenwei, Gao, Jianfei, Liu, Kuikun, Wang, Ziyi, Li, Shuaibin, Zhao, Qian, Huang, Haian, Cao, Weihan, Liu, Jiangning, Liu, Hongwei, Liu, Junnan, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
von: Shen, Junhao, et al.
Veröffentlicht: (2025)
von: Shen, Junhao, et al.
Veröffentlicht: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
Rectifying LLM Thought from Lens of Optimization
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
von: Zhao, Haiteng, et al.
Veröffentlicht: (2025)
von: Zhao, Haiteng, et al.
Veröffentlicht: (2025)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
von: Zhao, Yufeng, et al.
Veröffentlicht: (2025)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
von: Chen, Zehui, et al.
Veröffentlicht: (2024)
von: Chen, Zehui, et al.
Veröffentlicht: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
von: Zhang, Chuyu, et al.
Veröffentlicht: (2024)
von: Zhang, Chuyu, et al.
Veröffentlicht: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
Rethinking Verification for LLM Code Generation: From Generation to Testing
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
von: Chen, Zehui, et al.
Veröffentlicht: (2024)
von: Chen, Zehui, et al.
Veröffentlicht: (2024)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
von: Song, Zifan, et al.
Veröffentlicht: (2024)
von: Song, Zifan, et al.
Veröffentlicht: (2024)
Exploring the MBTI distribution among Chinese undergraduate physics students: the influence of family income on career trajectories
von: Bai, Songyang, et al.
Veröffentlicht: (2024)
von: Bai, Songyang, et al.
Veröffentlicht: (2024)
MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
von: Fang, Youqing, et al.
Veröffentlicht: (2026)
von: Fang, Youqing, et al.
Veröffentlicht: (2026)
Training Language Models to Critique With Multi-agent Feedback
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation
von: Gao, Songyang, et al.
Veröffentlicht: (2026)
von: Gao, Songyang, et al.
Veröffentlicht: (2026)
Learning Spatial Awareness for Laparoscopic Surgery with AI Assisted Visual Feedback
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
The adaptive EM schemes for McKean-Vlasov SDEs with common noise in finite and infinite horizons
von: Liu, Hu, et al.
Veröffentlicht: (2025)
von: Liu, Hu, et al.
Veröffentlicht: (2025)
Cross: A Delay Based Congestion Control Method for RTP Media
von: Zhang, Songyang, et al.
Veröffentlicht: (2024)
von: Zhang, Songyang, et al.
Veröffentlicht: (2024)
Fake Alignment: Are LLMs Really Aligned Well?
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
von: Chen, Sizhou, et al.
Veröffentlicht: (2023)
von: Chen, Sizhou, et al.
Veröffentlicht: (2023)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
von: Li, Mo, et al.
Veröffentlicht: (2024)
von: Li, Mo, et al.
Veröffentlicht: (2024)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
von: Zhang, Taolin, et al.
Veröffentlicht: (2025)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
Pre-Trained Policy Discriminators are General Reward Models
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025) -
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
von: Liu, Shudong, et al.
Veröffentlicht: (2025) -
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024) -
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
von: Shen, Junhao, et al.
Veröffentlicht: (2025) -
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)