Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gao, Songyang, Gu, Yuzhe, Wu, Zijian, Kong, Lingkai, Zhang, Wenwei, Cai, Zhongrui, Zheng, Fan, Ma, Tianyou, Shen, Junhao, Zhao, Haiteng, Zhang, Duanyang, Zhang, Huilun, Liu, Kuikun, Lyu, Chengqi, Duan, Yanhui, Chen, Chiyu, Ma, Ningsheng, Gao, Jianfei, Lyu, Han, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
par: Hua, Zhouqi, et autres
Publié: (2025)
par: Hua, Zhouqi, et autres
Publié: (2025)
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
par: Shen, Junhao, et autres
Publié: (2025)
par: Shen, Junhao, et autres
Publié: (2025)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
par: Zhao, Haiteng, et autres
Publié: (2025)
par: Zhao, Haiteng, et autres
Publié: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
par: Gu, Yuzhe, et autres
Publié: (2025)
par: Gu, Yuzhe, et autres
Publié: (2025)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
par: Wu, Zijian, et autres
Publié: (2025)
par: Wu, Zijian, et autres
Publié: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
par: Ji, Ziwei, et autres
Publié: (2024)
par: Ji, Ziwei, et autres
Publié: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
par: Gu, Yuzhe, et autres
Publié: (2024)
par: Gu, Yuzhe, et autres
Publié: (2024)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
par: Lyu, Chengqi, et autres
Publié: (2025)
par: Lyu, Chengqi, et autres
Publié: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
par: Liu, Shudong, et autres
Publié: (2025)
par: Liu, Shudong, et autres
Publié: (2025)
Are Your LLMs Capable of Stable Reasoning?
par: Liu, Junnan, et autres
Publié: (2024)
par: Liu, Junnan, et autres
Publié: (2024)
AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
par: Song, Zifan, et autres
Publié: (2024)
par: Song, Zifan, et autres
Publié: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
par: Zhang, Chuyu, et autres
Publié: (2024)
par: Zhang, Chuyu, et autres
Publié: (2024)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
par: Zhao, Zhonghan, et autres
Publié: (2025)
par: Zhao, Zhonghan, et autres
Publié: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
par: Chen, Zehui, et autres
Publié: (2023)
par: Chen, Zehui, et autres
Publié: (2023)
Fake Alignment: Are LLMs Really Aligned Well?
par: Wang, Yixu, et autres
Publié: (2023)
par: Wang, Yixu, et autres
Publié: (2023)
Training Language Models to Critique With Multi-agent Feedback
par: Lan, Tian, et autres
Publié: (2024)
par: Lan, Tian, et autres
Publié: (2024)
Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
par: Chen, Zehui, et autres
Publié: (2024)
par: Chen, Zehui, et autres
Publié: (2024)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
par: Chen, Zehui, et autres
Publié: (2024)
par: Chen, Zehui, et autres
Publié: (2024)
Rethinking Verification for LLM Code Generation: From Generation to Testing
par: Ma, Zihan, et autres
Publié: (2025)
par: Ma, Zihan, et autres
Publié: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
par: Li, Rongjie, et autres
Publié: (2024)
par: Li, Rongjie, et autres
Publié: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
par: Wang, Chonghua, et autres
Publié: (2024)
par: Wang, Chonghua, et autres
Publié: (2024)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
par: Li, Peiji, et autres
Publié: (2025)
par: Li, Peiji, et autres
Publié: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
par: Liu, Hongwei, et autres
Publié: (2024)
par: Liu, Hongwei, et autres
Publié: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
par: Ying, Huaiyuan, et autres
Publié: (2024)
par: Ying, Huaiyuan, et autres
Publié: (2024)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
par: Wu, Zijian, et autres
Publié: (2024)
par: Wu, Zijian, et autres
Publié: (2024)
Mastering Olympiad-Level Physics with Artificial Intelligence
par: Jian, Dong-Shan, et autres
Publié: (2025)
par: Jian, Dong-Shan, et autres
Publié: (2025)
Towards Imperceptible Adversarial Attacks for Time Series Classification with Local Perturbations and Frequency Analysis
par: Gu, Wenwei, et autres
Publié: (2025)
par: Gu, Wenwei, et autres
Publié: (2025)
Collaborative Performance Prediction for Large Language Models
par: Zhang, Qiyuan, et autres
Publié: (2024)
par: Zhang, Qiyuan, et autres
Publié: (2024)
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
par: He, Chaoqun, et autres
Publié: (2024)
par: He, Chaoqun, et autres
Publié: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
par: Zhuo, Jingming, et autres
Publié: (2024)
par: Zhuo, Jingming, et autres
Publié: (2024)
HUMAN RESOURCE DEVELOPMENT AND SOCIAL EMPOWERMENT: A HOLISTIC FRAMEWORK FOR SUSTAINABLE COMMUNITY GROWTH
par: Amiya Bhaumik, Lyu Wenwei, Nandar Win
Publié: (2026)
par: Amiya Bhaumik, Lyu Wenwei, Nandar Win
Publié: (2026)
The adaptive EM schemes for McKean-Vlasov SDEs with common noise in finite and infinite horizons
par: Liu, Hu, et autres
Publié: (2025)
par: Liu, Hu, et autres
Publié: (2025)
Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction
par: Su, Wuqi, et autres
Publié: (2026)
par: Su, Wuqi, et autres
Publié: (2026)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
par: Li, Zenan, et autres
Publié: (2025)
par: Li, Zenan, et autres
Publié: (2025)
Exploring the MBTI distribution among Chinese undergraduate physics students: the influence of family income on career trajectories
par: Bai, Songyang, et autres
Publié: (2024)
par: Bai, Songyang, et autres
Publié: (2024)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
par: Chen, Sizhou, et autres
Publié: (2023)
par: Chen, Sizhou, et autres
Publié: (2023)
InternLM-Law: An Open Source Chinese Legal Large Language Model
par: Fei, Zhiwei, et autres
Publié: (2024)
par: Fei, Zhiwei, et autres
Publié: (2024)
OpenCompass: A Universal Evaluation Platform for Large Language Models
par: Cao, Maosong, et autres
Publié: (2026)
par: Cao, Maosong, et autres
Publié: (2026)
Recent Progress of Low‐Dimensional Metal‐Organic Frameworks for Aqueous Zinc‐Based Batteries
par: Hanfang Xing, et autres
Publié: (2024)
par: Hanfang Xing, et autres
Publié: (2024)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
par: Gao, Bofei, et autres
Publié: (2024)
par: Gao, Bofei, et autres
Publié: (2024)
Documents similaires
-
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
par: Hua, Zhouqi, et autres
Publié: (2025) -
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
par: Shen, Junhao, et autres
Publié: (2025) -
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
par: Zhao, Haiteng, et autres
Publié: (2025) -
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
par: Gu, Yuzhe, et autres
Publié: (2025) -
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
par: Wu, Zijian, et autres
Publié: (2025)