Generalized Linear Markov Decision Process

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Sinian, Zhang, Kaicheng, Xu, Ziping, Cai, Tianxi, Zhou, Doudou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912408378802176
author Zhang, Sinian
Zhang, Kaicheng
Xu, Ziping
Cai, Tianxi
Zhou, Doudou
author_facet Zhang, Sinian
Zhang, Kaicheng
Xu, Ziping
Cai, Tianxi
Zhou, Doudou
contents The linear Markov Decision Process (MDP) framework offers a principled foundation for reinforcement learning (RL) with strong theoretical guarantees and sample efficiency. However, its restrictive assumption-that both transition dynamics and reward functions are linear in the same feature space-limits its applicability in real-world domains, where rewards often exhibit nonlinear or discrete structures. Motivated by applications such as healthcare and e-commerce, where data is scarce and reward signals can be binary or count-valued, we propose the Generalized Linear MDP (GLMDP) framework-an extension of the linear MDP framework-that models rewards using generalized linear models (GLMs) while maintaining linear transition dynamics. We establish the Bellman completeness of GLMDPs with respect to a new function class that accommodates nonlinear rewards and develop two offline RL algorithms: Generalized Pessimistic Value Iteration (GPEVI) and a semi-supervised variant (SS-GPEVI) that utilizes both labeled and unlabeled trajectories. Our algorithms achieve theoretical guarantees on policy suboptimality and demonstrate improved sample efficiency in settings where reward labels are expensive or limited.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalized Linear Markov Decision Process
Zhang, Sinian
Zhang, Kaicheng
Xu, Ziping
Cai, Tianxi
Zhou, Doudou
Machine Learning
The linear Markov Decision Process (MDP) framework offers a principled foundation for reinforcement learning (RL) with strong theoretical guarantees and sample efficiency. However, its restrictive assumption-that both transition dynamics and reward functions are linear in the same feature space-limits its applicability in real-world domains, where rewards often exhibit nonlinear or discrete structures. Motivated by applications such as healthcare and e-commerce, where data is scarce and reward signals can be binary or count-valued, we propose the Generalized Linear MDP (GLMDP) framework-an extension of the linear MDP framework-that models rewards using generalized linear models (GLMs) while maintaining linear transition dynamics. We establish the Bellman completeness of GLMDPs with respect to a new function class that accommodates nonlinear rewards and develop two offline RL algorithms: Generalized Pessimistic Value Iteration (GPEVI) and a semi-supervised variant (SS-GPEVI) that utilizes both labeled and unlabeled trajectories. Our algorithms achieve theoretical guarantees on policy suboptimality and demonstrate improved sample efficiency in settings where reward labels are expensive or limited.
title Generalized Linear Markov Decision Process
topic Machine Learning
url https://arxiv.org/abs/2506.00818