Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Fuente:
arXiv
Saved in:
| Main Authors: | Bereket, Michael, Leskovec, Jure |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TimeGraphs: Graph-based Temporal Reasoning
by: Maheshwari, Paridhi, et al.
Published: (2024)
by: Maheshwari, Paridhi, et al.
Published: (2024)
RelGNN: Composite Message Passing for Relational Deep Learning
by: Chen, Tianlang, et al.
Published: (2025)
by: Chen, Tianlang, et al.
Published: (2025)
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
by: Huang, Qian, et al.
Published: (2023)
by: Huang, Qian, et al.
Published: (2023)
Large Language Models are Good Relational Learners
by: Wu, Fang, et al.
Published: (2025)
by: Wu, Fang, et al.
Published: (2025)
Uncertainty Quantification for Forward and Inverse Problems of PDEs via Latent Global Evolution
by: Wu, Tailin, et al.
Published: (2024)
by: Wu, Tailin, et al.
Published: (2024)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Relational Deep Learning: Challenges, Foundations and Next-Generation Architectures
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
Learning over Positive and Negative Edges with Contrastive Message Passing
by: Pao-Huang, Peter, et al.
Published: (2026)
by: Pao-Huang, Peter, et al.
Published: (2026)
Modelling Cellular Perturbations with the Sparse Additive Mechanism Shift Variational Autoencoder
by: Bereket, Michael, et al.
Published: (2023)
by: Bereket, Michael, et al.
Published: (2023)
Agentic Uncertainty Reveals Agentic Overconfidence
by: Kaddour, Jean, et al.
Published: (2026)
by: Kaddour, Jean, et al.
Published: (2026)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
KumoRFM-2: Scaling Foundation Models for Relational Learning
by: Hudovernik, Valter, et al.
Published: (2026)
by: Hudovernik, Valter, et al.
Published: (2026)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
by: Kothapalli, Vignesh, et al.
Published: (2026)
by: Kothapalli, Vignesh, et al.
Published: (2026)
From Similarity to Superiority: Channel Clustering for Time Series Forecasting
by: Chen, Jialin, et al.
Published: (2024)
by: Chen, Jialin, et al.
Published: (2024)
Automated Hypothesis Validation with Agentic Sequential Falsifications
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
by: Lian, Yongsheng
Published: (2025)
by: Lian, Yongsheng
Published: (2025)
Mitigating Overconfidence in Out-of-Distribution Detection by Capturing Extreme Activations
by: Azizmalayeri, Mohammad, et al.
Published: (2024)
by: Azizmalayeri, Mohammad, et al.
Published: (2024)
Surface-based Molecular Design with Multi-modal Flow Matching
by: Wu, Fang, et al.
Published: (2026)
by: Wu, Fang, et al.
Published: (2026)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
TGM: a Modular and Efficient Library for Machine Learning on Temporal Graphs
by: Chmura, Jacob, et al.
Published: (2025)
by: Chmura, Jacob, et al.
Published: (2025)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
by: Wang, Jingyi, et al.
Published: (2026)
by: Wang, Jingyi, et al.
Published: (2026)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Relational Graph Transformer
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
by: Dwivedi, Vijay Prakash, et al.
Published: (2025)
Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases
by: Kocijan, Vid, et al.
Published: (2026)
by: Kocijan, Vid, et al.
Published: (2026)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Compositional Generative Inverse Design
by: Wu, Tailin, et al.
Published: (2024)
by: Wu, Tailin, et al.
Published: (2024)
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
by: Ranjan, Rishabh, et al.
Published: (2025)
by: Ranjan, Rishabh, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
by: Kumaran, Dharshan, et al.
Published: (2025)
by: Kumaran, Dharshan, et al.
Published: (2025)
Zero-shot causal learning
by: Nilforoshan, Hamed, et al.
Published: (2023)
by: Nilforoshan, Hamed, et al.
Published: (2023)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
by: Mansouri, Omar El, et al.
Published: (2025)
by: Mansouri, Omar El, et al.
Published: (2025)
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
by: Wu, Qingyuan, et al.
Published: (2026)
by: Wu, Qingyuan, et al.
Published: (2026)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
Physically Parameterized Differentiable MUSIC for DoA Estimation with Uncalibrated Arrays
by: Chatelier, Baptiste, et al.
Published: (2024)
by: Chatelier, Baptiste, et al.
Published: (2024)
Similar Items
-
TimeGraphs: Graph-based Temporal Reasoning
by: Maheshwari, Paridhi, et al.
Published: (2024) -
RelGNN: Composite Message Passing for Relational Deep Learning
by: Chen, Tianlang, et al.
Published: (2025) -
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
by: Huang, Qian, et al.
Published: (2023) -
Large Language Models are Good Relational Learners
by: Wu, Fang, et al.
Published: (2025) -
Uncertainty Quantification for Forward and Inverse Problems of PDEs via Latent Global Evolution
by: Wu, Tailin, et al.
Published: (2024)