Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Pengyi, Goncharova, Elizaveta, Kuznetsov, Andrey, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
by: Sevriugov, Egor, et al.
Published: (2024)
by: Sevriugov, Egor, et al.
Published: (2024)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
Sentence-Anchored Gist Compression for Long-Context LLMs
by: Tarasov, Dmitrii, et al.
Published: (2025)
by: Tarasov, Dmitrii, et al.
Published: (2025)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality
by: Chekalina, Viktoriia, et al.
Published: (2024)
by: Chekalina, Viktoriia, et al.
Published: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
by: Mezentsev, Gleb, et al.
Published: (2025)
by: Mezentsev, Gleb, et al.
Published: (2025)
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
by: Cheng, Zhoujun, et al.
Published: (2025)
by: Cheng, Zhoujun, et al.
Published: (2025)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
by: Huo, Yifu, et al.
Published: (2026)
by: Huo, Yifu, et al.
Published: (2026)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
by: Shan, Zikang, et al.
Published: (2026)
by: Shan, Zikang, et al.
Published: (2026)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025)
by: Bershatsky, Daniel, et al.
Published: (2025)
Outcome-based Exploration for LLM Reasoning
by: Song, Yuda, et al.
Published: (2025)
by: Song, Yuda, et al.
Published: (2025)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
by: Jin, Renren, et al.
Published: (2025)
by: Jin, Renren, et al.
Published: (2025)
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
by: Guo, Yiran, et al.
Published: (2026)
by: Guo, Yiran, et al.
Published: (2026)
Testing Uncertainty of Large Language Models for Physics Knowledge and Reasoning
by: Reganova, Elizaveta, et al.
Published: (2024)
by: Reganova, Elizaveta, et al.
Published: (2024)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Latent Feature Mining for Predictive Model Enhancement with Large Language Models
by: Li, Bingxuan, et al.
Published: (2024)
by: Li, Bingxuan, et al.
Published: (2024)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
by: Bao, Keqin, et al.
Published: (2025)
by: Bao, Keqin, et al.
Published: (2025)
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
by: Wan, Zhongwei, et al.
Published: (2026)
by: Wan, Zhongwei, et al.
Published: (2026)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
by: Zhou, Zhi, et al.
Published: (2025)
by: Zhou, Zhi, et al.
Published: (2025)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens
by: Dragunov, Nikita, et al.
Published: (2025)
by: Dragunov, Nikita, et al.
Published: (2025)
An Explainable Diagnostic Framework for Neurodegenerative Dementias via Reinforcement-Optimized LLM Reasoning
by: Zamai, Andrew, et al.
Published: (2025)
by: Zamai, Andrew, et al.
Published: (2025)
Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2025)
by: Zhu, Xinyu, et al.
Published: (2025)
ESQA: Event Sequences Question Answering
by: Abdullaeva, Irina, et al.
Published: (2024)
by: Abdullaeva, Irina, et al.
Published: (2024)
Similar Items
-
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023) -
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025) -
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024) -
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025) -
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
by: Sevriugov, Egor, et al.
Published: (2024)