Permutation Equivariant Model-based Offline Reinforcement Learning for Auto-bidding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mou, Zhiyu, Xu, Miao, Chen, Wei, Bai, Rongquan, Yu, Chuan, Xu, Jian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916805701795840
author Mou, Zhiyu
Xu, Miao
Chen, Wei
Bai, Rongquan
Yu, Chuan
Xu, Jian
author_facet Mou, Zhiyu
Xu, Miao
Chen, Wei
Bai, Rongquan
Yu, Chuan
Xu, Jian
contents Reinforcement learning (RL) for auto-bidding has shifted from using simplistic offline simulators (Simulation-based RL Bidding, SRLB) to offline RL on fixed real datasets (Offline RL Bidding, ORLB). However, ORLB policies are limited by the dataset's state space coverage, offering modest gains. While SRLB expands state coverage, its simulator-reality gap risks misleading policies. This paper introduces Model-based RL Bidding (MRLB), which learns an environment model from real data to bridge this gap. MRLB trains policies using both real and model-generated data, expanding state coverage beyond ORLB. To ensure model reliability, we propose: 1) A permutation equivariant model architecture for better generalization, and 2) A robust offline Q-learning method that pessimistically penalizes model errors. These form the Permutation Equivariant Model-based Offline RL (PE-MORL) algorithm. Real-world experiments show that PE-MORL outperforms state-of-the-art auto-bidding methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Permutation Equivariant Model-based Offline Reinforcement Learning for Auto-bidding
Mou, Zhiyu
Xu, Miao
Chen, Wei
Bai, Rongquan
Yu, Chuan
Xu, Jian
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) for auto-bidding has shifted from using simplistic offline simulators (Simulation-based RL Bidding, SRLB) to offline RL on fixed real datasets (Offline RL Bidding, ORLB). However, ORLB policies are limited by the dataset's state space coverage, offering modest gains. While SRLB expands state coverage, its simulator-reality gap risks misleading policies. This paper introduces Model-based RL Bidding (MRLB), which learns an environment model from real data to bridge this gap. MRLB trains policies using both real and model-generated data, expanding state coverage beyond ORLB. To ensure model reliability, we propose: 1) A permutation equivariant model architecture for better generalization, and 2) A robust offline Q-learning method that pessimistically penalizes model errors. These form the Permutation Equivariant Model-based Offline RL (PE-MORL) algorithm. Real-world experiments show that PE-MORL outperforms state-of-the-art auto-bidding methods.
title Permutation Equivariant Model-based Offline Reinforcement Learning for Auto-bidding
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.17919