Adaptive Robust Estimator for Multi-Agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhongyi, Tian, Wan, Chen, Jingyu, Huang, Kangyao, Zhang, Huiming, Yang, Hui, Ren, Tao, Jiang, Jinyang, Peng, Yijie, Ban, Yikun, Zhuang, Fuzhen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915882992664576
author Li, Zhongyi
Tian, Wan
Chen, Jingyu
Huang, Kangyao
Zhang, Huiming
Yang, Hui
Ren, Tao
Jiang, Jinyang
Peng, Yijie
Ban, Yikun
Zhuang, Fuzhen
author_facet Li, Zhongyi
Tian, Wan
Chen, Jingyu
Huang, Kangyao
Zhang, Huiming
Yang, Hui
Ren, Tao
Jiang, Jinyang
Peng, Yijie
Ban, Yikun
Zhuang, Fuzhen
contents Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that blurs generation, critique, and revision, making credit assignment across agents difficult. Moreover, policy optimization in this setting is vulnerable to heavy-tailed and noisy rewards, which can bias advantage estimation and trigger unstable or even divergent training. To address both issues, we propose a robust multi-agent reinforcement learning framework for collaborative reasoning, consisting of two components: Dual-Agent Answer-Critique-Rewrite (DACR) and an Adaptive Robust Estimator (ARE). DACR decomposes reasoning into a structured three-stage pipeline: answer, critique, and rewrite, while enabling explicit attribution of each agent's marginal contribution to its partner's performance. ARE provides robust estimation of batch experience means during multi-agent policy optimization. Across mathematical reasoning and embodied intelligence benchmarks, even under noisy rewards, our method consistently outperforms the baseline in both homogeneous and heterogeneous settings. These results indicate stronger robustness to reward noise and more stable training dynamics, effectively preventing optimization failures caused by noisy reward signals.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21574
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
Li, Zhongyi
Tian, Wan
Chen, Jingyu
Huang, Kangyao
Zhang, Huiming
Yang, Hui
Ren, Tao
Jiang, Jinyang
Peng, Yijie
Ban, Yikun
Zhuang, Fuzhen
Artificial Intelligence
Multi-agent collaboration has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models, yet it suffers from interaction-level ambiguity that blurs generation, critique, and revision, making credit assignment across agents difficult. Moreover, policy optimization in this setting is vulnerable to heavy-tailed and noisy rewards, which can bias advantage estimation and trigger unstable or even divergent training. To address both issues, we propose a robust multi-agent reinforcement learning framework for collaborative reasoning, consisting of two components: Dual-Agent Answer-Critique-Rewrite (DACR) and an Adaptive Robust Estimator (ARE). DACR decomposes reasoning into a structured three-stage pipeline: answer, critique, and rewrite, while enabling explicit attribution of each agent's marginal contribution to its partner's performance. ARE provides robust estimation of batch experience means during multi-agent policy optimization. Across mathematical reasoning and embodied intelligence benchmarks, even under noisy rewards, our method consistently outperforms the baseline in both homogeneous and heterogeneous settings. These results indicate stronger robustness to reward noise and more stable training dynamics, effectively preventing optimization failures caused by noisy reward signals.
title Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2603.21574