Efficient Duple Perturbation Robustness in Low-rank MDPs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Yang, Ma, Haitong, Dai, Bo, Li, Na
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929312111788032
author Hu, Yang
Ma, Haitong
Dai, Bo
Li, Na
author_facet Hu, Yang
Ma, Haitong
Dai, Bo
Li, Na
contents The pursuit of robustness has recently been a popular topic in reinforcement learning (RL) research, yet the existing methods generally suffer from efficiency issues that obstruct their real-world implementation. In this paper, we introduce duple perturbation robustness, i.e. perturbation on both the feature and factor vectors for low-rank Markov decision processes (MDPs), via a novel characterization of $(ξ,η)$-ambiguity sets. The novel robust MDP formulation is compatible with the function representation view, and therefore, is naturally applicable to practical RL problems with large or even continuous state-action spaces. Meanwhile, it also gives rise to a provably efficient and practical algorithm with theoretical convergence rate guarantee. Examples are designed to justify the new robustness concept, and algorithmic efficiency is supported by both theoretical bounds and numerical simulations.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08089
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Duple Perturbation Robustness in Low-rank MDPs
Hu, Yang
Ma, Haitong
Dai, Bo
Li, Na
Machine Learning
Optimization and Control
The pursuit of robustness has recently been a popular topic in reinforcement learning (RL) research, yet the existing methods generally suffer from efficiency issues that obstruct their real-world implementation. In this paper, we introduce duple perturbation robustness, i.e. perturbation on both the feature and factor vectors for low-rank Markov decision processes (MDPs), via a novel characterization of $(ξ,η)$-ambiguity sets. The novel robust MDP formulation is compatible with the function representation view, and therefore, is naturally applicable to practical RL problems with large or even continuous state-action spaces. Meanwhile, it also gives rise to a provably efficient and practical algorithm with theoretical convergence rate guarantee. Examples are designed to justify the new robustness concept, and algorithmic efficiency is supported by both theoretical bounds and numerical simulations.
title Efficient Duple Perturbation Robustness in Low-rank MDPs
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2404.08089