MaRCA: Multi-Agent Reinforcement Learning for Dynamic Computation Allocation in Large-Scale Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Wan, Zang, Xinyi, Zhao, Yudong, Zou, Yusi, Lu, Yunfei, Tong, Junbo, Liu, Yang, Li, Ming, Shi, Jiani, Yang, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918266237091840
author Jiang, Wan
Zang, Xinyi
Zhao, Yudong
Zou, Yusi
Lu, Yunfei
Tong, Junbo
Liu, Yang
Li, Ming
Shi, Jiani
Yang, Xin
author_facet Jiang, Wan
Zang, Xinyi
Zhao, Yudong
Zou, Yusi
Lu, Yunfei
Tong, Junbo
Liu, Yang
Li, Ming
Shi, Jiani
Yang, Xin
contents Modern recommender systems face significant computational challenges due to growing model complexity and traffic scale, making efficient computation allocation critical for maximizing business revenue. Existing approaches typically simplify multi-stage computation resource allocation, neglecting inter-stage dependencies, thus limiting global optimality. In this paper, we propose MaRCA, a multi-agent reinforcement learning framework for end-to-end computation resource allocation in large-scale recommender systems. MaRCA models the stages of a recommender system as cooperative agents, using Centralized Training with Decentralized Execution (CTDE) to optimize revenue under computation resource constraints. We introduce an AutoBucket TestBench for accurate computation cost estimation, and a Model Predictive Control (MPC)-based Revenue-Cost Balancer to proactively forecast traffic loads and adjust the revenue-cost trade-off accordingly. Since its end-to-end deployment in the advertising pipeline of a leading global e-commerce platform in November 2024, MaRCA has consistently handled hundreds of billions of ad requests per day and has delivered a 16.67% revenue uplift using existing computation resources.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MaRCA: Multi-Agent Reinforcement Learning for Dynamic Computation Allocation in Large-Scale Recommender Systems
Jiang, Wan
Zang, Xinyi
Zhao, Yudong
Zou, Yusi
Lu, Yunfei
Tong, Junbo
Liu, Yang
Li, Ming
Shi, Jiani
Yang, Xin
Information Retrieval
Machine Learning
Multiagent Systems
H.3.3; I.2.11
Modern recommender systems face significant computational challenges due to growing model complexity and traffic scale, making efficient computation allocation critical for maximizing business revenue. Existing approaches typically simplify multi-stage computation resource allocation, neglecting inter-stage dependencies, thus limiting global optimality. In this paper, we propose MaRCA, a multi-agent reinforcement learning framework for end-to-end computation resource allocation in large-scale recommender systems. MaRCA models the stages of a recommender system as cooperative agents, using Centralized Training with Decentralized Execution (CTDE) to optimize revenue under computation resource constraints. We introduce an AutoBucket TestBench for accurate computation cost estimation, and a Model Predictive Control (MPC)-based Revenue-Cost Balancer to proactively forecast traffic loads and adjust the revenue-cost trade-off accordingly. Since its end-to-end deployment in the advertising pipeline of a leading global e-commerce platform in November 2024, MaRCA has consistently handled hundreds of billions of ad requests per day and has delivered a 16.67% revenue uplift using existing computation resources.
title MaRCA: Multi-Agent Reinforcement Learning for Dynamic Computation Allocation in Large-Scale Recommender Systems
topic Information Retrieval
Machine Learning
Multiagent Systems
H.3.3; I.2.11
url https://arxiv.org/abs/2512.24325