Outcome-based Exploration for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yuda, Kempe, Julia, Munos, Remi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025)
by: Arnal, Charles, et al.
Published: (2025)
Accelerating Unbiased LLM Evaluation via Synthetic Feedback
by: Zhou, Zhaoyi, et al.
Published: (2025)
by: Zhou, Zhaoyi, et al.
Published: (2025)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
by: Sundaram, Shobhita, et al.
Published: (2026)
by: Sundaram, Shobhita, et al.
Published: (2026)
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
by: Wan, Zhongwei, et al.
Published: (2026)
by: Wan, Zhongwei, et al.
Published: (2026)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
by: Li, Pengyi, et al.
Published: (2026)
by: Li, Pengyi, et al.
Published: (2026)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
Soft Tokens, Hard Truths
by: Butt, Natasha, et al.
Published: (2025)
by: Butt, Natasha, et al.
Published: (2025)
A Tale of Tails: Model Collapse as a Change of Scaling Laws
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
by: Labiad, Ismail, et al.
Published: (2025)
by: Labiad, Ismail, et al.
Published: (2025)
Knowledge Graph-Assisted LLM Post-Training for Enhanced Legal Reasoning
by: Song, Dezhao, et al.
Published: (2026)
by: Song, Dezhao, et al.
Published: (2026)
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Iteration Head: A Mechanistic Study of Chain-of-Thought
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
by: Mu, Yongyu, et al.
Published: (2026)
by: Mu, Yongyu, et al.
Published: (2026)
Nudging the Boundaries of LLM Reasoning
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
by: Lyu, Chengqi, et al.
Published: (2025)
by: Lyu, Chengqi, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Efficiently Scaling LLM Reasoning with Certaindex
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark
by: Yuan, Zhangdie, et al.
Published: (2025)
by: Yuan, Zhangdie, et al.
Published: (2025)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
by: Song, Yuda, et al.
Published: (2024)
by: Song, Yuda, et al.
Published: (2024)
Multi-LLM QA with Embodied Exploration
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
On a few pitfalls in KL divergence gradient estimation for RL
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Lexical Hints of Accuracy in LLM Reasoning Chains
by: Vanhoyweghen, Arne, et al.
Published: (2025)
by: Vanhoyweghen, Arne, et al.
Published: (2025)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2025)
by: Zhu, Xinyu, et al.
Published: (2025)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
by: You, Weiqiu, et al.
Published: (2025)
by: You, Weiqiu, et al.
Published: (2025)
AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
by: He, Lixuan, et al.
Published: (2025)
by: He, Lixuan, et al.
Published: (2025)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
by: Li, Ziheng, et al.
Published: (2026)
by: Li, Ziheng, et al.
Published: (2026)
Graphical Reasoning: LLM-based Semi-Open Relation Extraction
by: Tao, Yicheng, et al.
Published: (2024)
by: Tao, Yicheng, et al.
Published: (2024)
From Belief Entrenchment to Robust Reasoning in LLM Agents
by: Oh, Jihwan, et al.
Published: (2025)
by: Oh, Jihwan, et al.
Published: (2025)
Process Supervision of Confidence Margin for Calibrated LLM Reasoning
by: Wang, Liaoyaqi, et al.
Published: (2026)
by: Wang, Liaoyaqi, et al.
Published: (2026)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
Similar Items
-
Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
by: Arnal, Charles, et al.
Published: (2025) -
Accelerating Unbiased LLM Evaluation via Synthetic Feedback
by: Zhou, Zhaoyi, et al.
Published: (2025) -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
by: Sundaram, Shobhita, et al.
Published: (2026) -
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026) -
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)