ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ai, Rui, Pan, Yu, Simchi-Levi, David, Wang, Chonghuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
von: Ai, Rui, et al.
Veröffentlicht: (2025)
von: Ai, Rui, et al.
Veröffentlicht: (2025)
LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency
von: Li, Jiachun, et al.
Veröffentlicht: (2026)
von: Li, Jiachun, et al.
Veröffentlicht: (2026)
ShapShift: Explaining Model Prediction Shifts with Subgroup Conditional Shapley Values
von: Bewley, Tom, et al.
Veröffentlicht: (2026)
von: Bewley, Tom, et al.
Veröffentlicht: (2026)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
von: Akkas, Selahattin, et al.
Veröffentlicht: (2025)
von: Akkas, Selahattin, et al.
Veröffentlicht: (2025)
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
von: Chen, Shuze, et al.
Veröffentlicht: (2024)
von: Chen, Shuze, et al.
Veröffentlicht: (2024)
ShapG: new feature importance method based on the Shapley value
von: Zhao, Chi, et al.
Veröffentlicht: (2024)
von: Zhao, Chi, et al.
Veröffentlicht: (2024)
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
Large Language Models for Supply Chain Decisions
von: Simchi-Levi, David, et al.
Veröffentlicht: (2025)
von: Simchi-Levi, David, et al.
Veröffentlicht: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
von: Wang, Zhijie
Veröffentlicht: (2026)
von: Wang, Zhijie
Veröffentlicht: (2026)
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
What Matters in Data for DPO?
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
GRPO is Secretly a Process Reward Model
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
When Right Meets Wrong: Bilateral Context Conditioning with Reward-Confidence Correction for GRPO
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Multi-agent Adaptive Mechanism Design
von: Han, Qiushi, et al.
Veröffentlicht: (2025)
von: Han, Qiushi, et al.
Veröffentlicht: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
von: Li, Yanming, et al.
Veröffentlicht: (2026)
von: Li, Yanming, et al.
Veröffentlicht: (2026)
Owen-based Semantics and Hierarchy-Aware Explanation (O-Shap)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2026)
Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
von: Yang, Pei, et al.
Veröffentlicht: (2025)
von: Yang, Pei, et al.
Veröffentlicht: (2025)
Improving LLM-Generated Code Quality with GRPO
von: Robeyns, Maxime, et al.
Veröffentlicht: (2025)
von: Robeyns, Maxime, et al.
Veröffentlicht: (2025)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
von: Mansouri, Omar El, et al.
Veröffentlicht: (2025)
von: Mansouri, Omar El, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
GESA: Graph-Enhanced Semantic Allocation for Generalized, Fair, and Explainable Candidate-Role Matching
von: Shah, Rishi Ashish, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Ashish, et al.
Veröffentlicht: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2026)
von: Liu, Shuze Daniel, et al.
Veröffentlicht: (2026)
iGRPO: Self-Feedback-Driven LLM Reasoning
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution
von: Feng, Xiachong, et al.
Veröffentlicht: (2026)
von: Feng, Xiachong, et al.
Veröffentlicht: (2026)
SyntaxShap: Syntax-aware Explainability Method for Text Generation
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
von: Amara, Kenza, et al.
Veröffentlicht: (2024)
Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
Designing Service Systems from Textual Evidence
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
von: Huang, Runhui, et al.
Veröffentlicht: (2026)
von: Huang, Runhui, et al.
Veröffentlicht: (2026)
Fast-DataShapley: Neural Modeling for Training Data Valuation
von: Sun, Haifeng, et al.
Veröffentlicht: (2025)
von: Sun, Haifeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
von: Ai, Rui, et al.
Veröffentlicht: (2025) -
LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency
von: Li, Jiachun, et al.
Veröffentlicht: (2026) -
ShapShift: Explaining Model Prediction Shifts with Subgroup Conditional Shapley Values
von: Bewley, Tom, et al.
Veröffentlicht: (2026) -
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026) -
DistShap: Scalable GNN Explanations with Distributed Shapley Values
von: Akkas, Selahattin, et al.
Veröffentlicht: (2025)