Graph-Attentive MAPPO for Dynamic Retail Pricing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908622955479040
author Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
author_facet Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
contents Dynamic pricing in retail requires policies that adapt to shifting demand while coordinating decisions across related products. We present a systematic empirical study of multi-agent reinforcement learning for retail price optimization, comparing a strong MAPPO baseline with a graph-attention-augmented variant (MAPPO+GAT) that leverages learned interactions among products. Using a simulated pricing environment derived from real transaction data, we evaluate profit, stability across random seeds, fairness across products, and training efficiency under a standardized evaluation protocol. The results indicate that MAPPO provides a robust and reproducible foundation for portfolio-level price control, and that MAPPO+GAT further enhances performance by sharing information over the product graph without inducing excessive price volatility. These results indicate that graph-integrated MARL provides a more scalable and stable solution than independent learners for dynamic retail pricing, offering practical advantages in multi-product decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Graph-Attentive MAPPO for Dynamic Retail Pricing
Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Artificial Intelligence
Machine Learning
Dynamic pricing in retail requires policies that adapt to shifting demand while coordinating decisions across related products. We present a systematic empirical study of multi-agent reinforcement learning for retail price optimization, comparing a strong MAPPO baseline with a graph-attention-augmented variant (MAPPO+GAT) that leverages learned interactions among products. Using a simulated pricing environment derived from real transaction data, we evaluate profit, stability across random seeds, fairness across products, and training efficiency under a standardized evaluation protocol. The results indicate that MAPPO provides a robust and reproducible foundation for portfolio-level price control, and that MAPPO+GAT further enhances performance by sharing information over the product graph without inducing excessive price volatility. These results indicate that graph-integrated MARL provides a more scalable and stable solution than independent learners for dynamic retail pricing, offering practical advantages in multi-product decision-making.
title Graph-Attentive MAPPO for Dynamic Retail Pricing
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.00039