Performative Policy Gradient: Optimality in Performative Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Basu, Debabrota, Das, Udvas, Driss, Brahim, Mukherjee, Uddalak
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917240478105600
author Basu, Debabrota
Das, Udvas
Driss, Brahim
Mukherjee, Uddalak
author_facet Basu, Debabrota
Das, Udvas
Driss, Brahim
Mukherjee, Uddalak
contents Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and further introduce the Performative Policy Gradient algorithm (PePG). PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Furthermore, our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20576
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performative Policy Gradient: Optimality in Performative Reinforcement Learning
Basu, Debabrota
Das, Udvas
Driss, Brahim
Mukherjee, Uddalak
Machine Learning
Artificial Intelligence
Optimization and Control
Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learning, the RL counterpart remains under-explored. In this paper, we prove the performative counterparts of the performance difference lemma and the policy gradient theorem in RL, and further introduce the Performative Policy Gradient algorithm (PePG). PePG is the first policy gradient algorithm designed to account for performativity in RL. Under softmax parametrisation, and also with and without entropy regularisation, we prove that PePG converges to performatively optimal policies, i.e. policies that remain optimal under the distribution shifts induced by themselves. Thus, PePG significantly extends the prior works in Performative RL that achieves performative stability but not optimality. Furthermore, our empirical analysis on standard performative RL environments validate that PePG outperforms the existing performative RL algorithms aiming for stability.
title Performative Policy Gradient: Optimality in Performative Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2512.20576