Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ang, Yuan, Zhihang, Zhang, Yang, Liu, Shouda, Wang, Yisen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation
by: Di Scala, Daan, et al.
Published: (2026)
by: Di Scala, Daan, et al.
Published: (2026)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
PID: Prompt-Independent Data Protection Against Latent Diffusion Models
by: Li, Ang, et al.
Published: (2024)
by: Li, Ang, et al.
Published: (2024)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
CaRT: Teaching LLM Agents to Know When They Know Enough
by: Liu, Grace, et al.
Published: (2025)
by: Liu, Grace, et al.
Published: (2025)
The Confidence Paradox: Can LLM Know When It's Wrong
by: Tripathi, Sahil, et al.
Published: (2025)
by: Tripathi, Sahil, et al.
Published: (2025)
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
by: Ding, Chenlu, et al.
Published: (2026)
by: Ding, Chenlu, et al.
Published: (2026)
SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling
by: Guo, Yi, et al.
Published: (2025)
by: Guo, Yi, et al.
Published: (2025)
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025)
by: Kong, Deyang, et al.
Published: (2025)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Application of LLM Guided Reinforcement Learning in Formation Control with Collision Avoidance
by: Yao, Chenhao, et al.
Published: (2025)
by: Yao, Chenhao, et al.
Published: (2025)
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
by: Zhang, Ru, et al.
Published: (2026)
by: Zhang, Ru, et al.
Published: (2026)
Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents
by: Mi, Qirui, et al.
Published: (2026)
by: Mi, Qirui, et al.
Published: (2026)
KnowPC: Knowledge-Driven Programmatic Reinforcement Learning for Zero-shot Coordination
by: Gu, Yin, et al.
Published: (2024)
by: Gu, Yin, et al.
Published: (2024)
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
by: Hao, Chenjie, et al.
Published: (2026)
by: Hao, Chenjie, et al.
Published: (2026)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
by: Zhang, Enci, et al.
Published: (2025)
by: Zhang, Enci, et al.
Published: (2025)
Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement Learning
by: Mu, Enhong, et al.
Published: (2025)
by: Mu, Enhong, et al.
Published: (2025)
PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach
by: Lin, Zhihao, et al.
Published: (2024)
by: Lin, Zhihao, et al.
Published: (2024)
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
by: Yu, Linhao, et al.
Published: (2026)
by: Yu, Linhao, et al.
Published: (2026)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training
by: Kang, Linjia, et al.
Published: (2026)
by: Kang, Linjia, et al.
Published: (2026)
Epistemic Artificial Intelligence is Essential for Machine Learning Models to Truly 'Know When They Do Not Know'
by: Manchingal, Shireen Kudukkil, et al.
Published: (2025)
by: Manchingal, Shireen Kudukkil, et al.
Published: (2025)
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
by: Duan, Jinhao, et al.
Published: (2025)
by: Duan, Jinhao, et al.
Published: (2025)
Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know
by: Manchingal, Shireen Kudukkil
Published: (2025)
by: Manchingal, Shireen Kudukkil
Published: (2025)
CVeDRL: An Efficient Code Verifier via Difficulty-aware Reinforcement Learning
by: Shi, Ji, et al.
Published: (2026)
by: Shi, Ji, et al.
Published: (2026)
Disentangle-then-Refine: LLM-Guided Decoupling and Structure-Aware Refinement for Graph Contrastive Learning
by: Li, Zhaoxing, et al.
Published: (2026)
by: Li, Zhaoxing, et al.
Published: (2026)
An Augmentation Overlap Theory of Contrastive Learning
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
Personalized Dynamic Difficulty Adjustment -- Imitation Learning Meets Reinforcement Learning
by: Fuchs, Ronja, et al.
Published: (2024)
by: Fuchs, Ronja, et al.
Published: (2024)
KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
by: Ren, Baochang, et al.
Published: (2025)
by: Ren, Baochang, et al.
Published: (2025)
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
by: Liu, Jun, et al.
Published: (2026)
by: Liu, Jun, et al.
Published: (2026)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
by: Huang, Zixuan, et al.
Published: (2026)
by: Huang, Zixuan, et al.
Published: (2026)
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
by: Li, Xuchen, et al.
Published: (2026)
by: Li, Xuchen, et al.
Published: (2026)
Similar Items
-
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
by: Li, Ang, et al.
Published: (2025) -
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025) -
"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation
by: Di Scala, Daan, et al.
Published: (2026) -
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
by: Li, Ang, et al.
Published: (2025) -
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)