Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Chengqi, Gao, Songyang, Gu, Yuzhe, Zhang, Wenwei, Gao, Jianfei, Liu, Kuikun, Wang, Ziyi, Li, Shuaibin, Zhao, Qian, Huang, Haian, Cao, Weihan, Liu, Jiangning, Liu, Hongwei, Liu, Junnan, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
by: Hua, Zhouqi, et al.
Published: (2025)
by: Hua, Zhouqi, et al.
Published: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
by: Liu, Shudong, et al.
Published: (2025)
by: Liu, Shudong, et al.
Published: (2025)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
by: Shen, Junhao, et al.
Published: (2025)
by: Shen, Junhao, et al.
Published: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
by: Gao, Songyang, et al.
Published: (2025)
by: Gao, Songyang, et al.
Published: (2025)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
by: Zhao, Zhonghan, et al.
Published: (2025)
by: Zhao, Zhonghan, et al.
Published: (2025)
Rectifying LLM Thought from Lens of Optimization
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
by: Gu, Yuzhe, et al.
Published: (2025)
by: Gu, Yuzhe, et al.
Published: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
by: Chen, Zehui, et al.
Published: (2023)
by: Chen, Zehui, et al.
Published: (2023)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
by: Zhao, Haiteng, et al.
Published: (2025)
by: Zhao, Haiteng, et al.
Published: (2025)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025)
by: Zhao, Yufeng, et al.
Published: (2025)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
by: Gu, Yuzhe, et al.
Published: (2024)
by: Gu, Yuzhe, et al.
Published: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
by: Chen, Zehui, et al.
Published: (2024)
by: Chen, Zehui, et al.
Published: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
by: Zhang, Chuyu, et al.
Published: (2024)
by: Zhang, Chuyu, et al.
Published: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
by: Ying, Huaiyuan, et al.
Published: (2024)
by: Ying, Huaiyuan, et al.
Published: (2024)
Rethinking Verification for LLM Code Generation: From Generation to Testing
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
by: Chen, Zehui, et al.
Published: (2024)
by: Chen, Zehui, et al.
Published: (2024)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
by: Song, Zifan, et al.
Published: (2024)
by: Song, Zifan, et al.
Published: (2024)
Exploring the MBTI distribution among Chinese undergraduate physics students: the influence of family income on career trajectories
by: Bai, Songyang, et al.
Published: (2024)
by: Bai, Songyang, et al.
Published: (2024)
MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
by: Fang, Youqing, et al.
Published: (2026)
by: Fang, Youqing, et al.
Published: (2026)
Training Language Models to Critique With Multi-agent Feedback
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation
by: Gao, Songyang, et al.
Published: (2026)
by: Gao, Songyang, et al.
Published: (2026)
Learning Spatial Awareness for Laparoscopic Surgery with AI Assisted Visual Feedback
by: Liu, Songyang, et al.
Published: (2025)
by: Liu, Songyang, et al.
Published: (2025)
The adaptive EM schemes for McKean-Vlasov SDEs with common noise in finite and infinite horizons
by: Liu, Hu, et al.
Published: (2025)
by: Liu, Hu, et al.
Published: (2025)
Cross: A Delay Based Congestion Control Method for RTP Media
by: Zhang, Songyang, et al.
Published: (2024)
by: Zhang, Songyang, et al.
Published: (2024)
Fake Alignment: Are LLMs Really Aligned Well?
by: Wang, Yixu, et al.
Published: (2023)
by: Wang, Yixu, et al.
Published: (2023)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
by: Cao, Maosong, et al.
Published: (2024)
by: Cao, Maosong, et al.
Published: (2024)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
by: Chen, Sizhou, et al.
Published: (2023)
by: Chen, Sizhou, et al.
Published: (2023)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
by: Li, Mo, et al.
Published: (2024)
by: Li, Mo, et al.
Published: (2024)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
by: Zhuo, Jingming, et al.
Published: (2024)
by: Zhuo, Jingming, et al.
Published: (2024)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Similar Items
-
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
by: Hua, Zhouqi, et al.
Published: (2025) -
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
by: Liu, Shudong, et al.
Published: (2025) -
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024) -
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
by: Shen, Junhao, et al.
Published: (2025) -
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
by: Gao, Songyang, et al.
Published: (2025)