Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yanming, Zhang, Xuelin, Lu, WenJie, Tang, Ziye, Wu, Maodong, Luo, Haotian, Wu, Tongtong, Peng, Zijie, Mi, Hongze, Feng, Yibo, Tan, Naiqiang, Huang, Chao, Chen, Hong, Shen, Li
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911433931882496
author Li, Yanming
Zhang, Xuelin
Lu, WenJie
Tang, Ziye
Wu, Maodong
Luo, Haotian
Wu, Tongtong
Peng, Zijie
Mi, Hongze
Feng, Yibo
Tan, Naiqiang
Huang, Chao
Chen, Hong
Shen, Li
author_facet Li, Yanming
Zhang, Xuelin
Lu, WenJie
Tang, Ziye
Wu, Maodong
Luo, Haotian
Wu, Tongtong
Peng, Zijie
Mi, Hongze
Feng, Yibo
Tan, Naiqiang
Huang, Chao
Chen, Hong
Shen, Li
contents Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems. However, training these systems remains notoriously difficult due to the credit assignment challenge, as it is often unclear which specific functional agent is responsible for the success or failure of decision trajectories. Existing methods typically rely on sparse or globally broadcast rewards, failing to capture individual contributions and leading to inefficient reinforcement learning. To address these limitations, we introduce the Shapley-based Hierarchical Attribution for Reinforcement Policy (SHARP), a novel framework for optimizing multi-agent reinforcement learning via precise credit attribution. SHARP effectively stabilizes training by normalizing agent-specific advantages across trajectory groups, primarily through a decomposed reward mechanism comprising a global broadcast-accuracy reward, a Shapley-based marginal-credit reward for each agent, and a tool-process reward to improve execution efficiency. Extensive experiments across various real-world benchmarks demonstrate that SHARP significantly outperforms recent state-of-the-art baselines, achieving average match improvements of 23.66% and 14.05% over single-agent and multi-agent approaches, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08335
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
Li, Yanming
Zhang, Xuelin
Lu, WenJie
Tang, Ziye
Wu, Maodong
Luo, Haotian
Wu, Tongtong
Peng, Zijie
Mi, Hongze
Feng, Yibo
Tan, Naiqiang
Huang, Chao
Chen, Hong
Shen, Li
Artificial Intelligence
Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems. However, training these systems remains notoriously difficult due to the credit assignment challenge, as it is often unclear which specific functional agent is responsible for the success or failure of decision trajectories. Existing methods typically rely on sparse or globally broadcast rewards, failing to capture individual contributions and leading to inefficient reinforcement learning. To address these limitations, we introduce the Shapley-based Hierarchical Attribution for Reinforcement Policy (SHARP), a novel framework for optimizing multi-agent reinforcement learning via precise credit attribution. SHARP effectively stabilizes training by normalizing agent-specific advantages across trajectory groups, primarily through a decomposed reward mechanism comprising a global broadcast-accuracy reward, a Shapley-based marginal-credit reward for each agent, and a tool-process reward to improve execution efficiency. Extensive experiments across various real-world benchmarks demonstrate that SHARP significantly outperforms recent state-of-the-art baselines, achieving average match improvements of 23.66% and 14.05% over single-agent and multi-agent approaches, respectively.
title Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
topic Artificial Intelligence
url https://arxiv.org/abs/2602.08335