Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhipeng, Qin, Xiaobo, Wu, Youbin, Ling, Yue, Ye, Qinghao, Zhao, Wayne Xin, Shi, Guang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
Reasoning with Exploration: An Entropy Perspective
von: Cheng, Daixuan, et al.
Veröffentlicht: (2025)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2025)
Disentangling Exploration of Large Language Models by Optimal Exploitation
von: Grams, Tim, et al.
Veröffentlicht: (2025)
von: Grams, Tim, et al.
Veröffentlicht: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
von: Peng, Chunyi, et al.
Veröffentlicht: (2025)
von: Peng, Chunyi, et al.
Veröffentlicht: (2025)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
von: Gao, Changjiang, et al.
Veröffentlicht: (2026)
von: Gao, Changjiang, et al.
Veröffentlicht: (2026)
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
von: Du, Yifan, et al.
Veröffentlicht: (2025)
von: Du, Yifan, et al.
Veröffentlicht: (2025)
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
ForesightKV: Optimizing KV Cache Eviction for Reasoning Models by Learning Long-Term Contribution
von: Dong, Zican, et al.
Veröffentlicht: (2026)
von: Dong, Zican, et al.
Veröffentlicht: (2026)
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data
von: Sun, Haoran, et al.
Veröffentlicht: (2024)
von: Sun, Haoran, et al.
Veröffentlicht: (2024)
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2025)
LLM Reasoning Engine: Specialized Training for Enhanced Mathematical Reasoning
von: Chen, Shuguang, et al.
Veröffentlicht: (2024)
von: Chen, Shuguang, et al.
Veröffentlicht: (2024)
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
Scale-Adaptive Balancing of Exploration and Exploitation in Classical Planning
von: Wissow, Stephen, et al.
Veröffentlicht: (2023)
von: Wissow, Stephen, et al.
Veröffentlicht: (2023)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
von: Chen, Qianben, et al.
Veröffentlicht: (2025)
von: Chen, Qianben, et al.
Veröffentlicht: (2025)
From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
von: Deng, Jia, et al.
Veröffentlicht: (2025)
von: Deng, Jia, et al.
Veröffentlicht: (2025)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
von: Peng, Jingyu, et al.
Veröffentlicht: (2025)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
HCR-Reasoner: Synergizing Large Language Models and Theory for Human-like Causal Reasoning
von: Zhang, Yanxi, et al.
Veröffentlicht: (2025)
von: Zhang, Yanxi, et al.
Veröffentlicht: (2025)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
von: Deng, Yimin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026) -
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025) -
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024) -
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025) -
Reasoning with Exploration: An Entropy Perspective
von: Cheng, Daixuan, et al.
Veröffentlicht: (2025)