MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Cuiling, Gan, Yaozhong, Xing, Junliang, Fu, Ying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908735190859776
author Wu, Cuiling
Gan, Yaozhong
Xing, Junliang
Fu, Ying
author_facet Wu, Cuiling
Gan, Yaozhong
Xing, Junliang
Fu, Ying
contents We propose Multi Agent Reflective Policy Optimization (MARPO) to alleviate the issue of sample inefficiency in multi agent reinforcement learning. MARPO consists of two key components: a reflection mechanism that leverages subsequent trajectories to enhance sample efficiency, and an asymmetric clipping mechanism that is derived from the KL divergence and dynamically adjusts the clipping range to improve training stability. We evaluate MARPO in classic multi agent environments, where it consistently outperforms other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning
Wu, Cuiling
Gan, Yaozhong
Xing, Junliang
Fu, Ying
Multiagent Systems
We propose Multi Agent Reflective Policy Optimization (MARPO) to alleviate the issue of sample inefficiency in multi agent reinforcement learning. MARPO consists of two key components: a reflection mechanism that leverages subsequent trajectories to enhance sample efficiency, and an asymmetric clipping mechanism that is derived from the KL divergence and dynamically adjusts the clipping range to improve training stability. We evaluate MARPO in classic multi agent environments, where it consistently outperforms other methods.
title MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning
topic Multiagent Systems
url https://arxiv.org/abs/2512.22832