SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shuai, Fernando, Heshan Devaka, Liu, Miao, Murugesan, Keerthiram, Lu, Songtao, Chen, Pin-Yu, Chen, Tianyi, Wang, Meng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910615527751680
author Zhang, Shuai
Fernando, Heshan Devaka
Liu, Miao
Murugesan, Keerthiram
Lu, Songtao
Chen, Pin-Yu
Chen, Tianyi
Wang, Meng
author_facet Zhang, Shuai
Fernando, Heshan Devaka
Liu, Miao
Murugesan, Keerthiram
Lu, Songtao
Chen, Pin-Yu
Chen, Tianyi
Wang, Meng
contents This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward mapping: the former characterizes the transition dynamics, and the latter characterizes the task-specific reward function. This Q-function decomposition, coupled with a policy improvement operator known as generalized policy improvement (GPI), reduces the sample complexity of finding the optimal Q-function, and thus the SF \& GPI framework exhibits promising empirical performance compared to traditional RL methods like Q-learning. However, its theoretical foundations remain largely unestablished, especially when learning the successor features using deep neural networks (SF-DQN). This paper studies the provable knowledge transfer using SFs-DQN in transfer RL problems. We establish the first convergence analysis with provable generalization guarantees for SF-DQN with GPI. The theory reveals that SF-DQN with GPI outperforms conventional RL approaches, such as deep Q-network, in terms of both faster convergence rate and better generalization. Numerical experiments on real and synthetic RL tasks support the superior performance of SF-DQN \& GPI, aligning with our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15920
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
Zhang, Shuai
Fernando, Heshan Devaka
Liu, Miao
Murugesan, Keerthiram
Lu, Songtao
Chen, Pin-Yu
Chen, Tianyi
Wang, Meng
Machine Learning
This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward mapping: the former characterizes the transition dynamics, and the latter characterizes the task-specific reward function. This Q-function decomposition, coupled with a policy improvement operator known as generalized policy improvement (GPI), reduces the sample complexity of finding the optimal Q-function, and thus the SF \& GPI framework exhibits promising empirical performance compared to traditional RL methods like Q-learning. However, its theoretical foundations remain largely unestablished, especially when learning the successor features using deep neural networks (SF-DQN). This paper studies the provable knowledge transfer using SFs-DQN in transfer RL problems. We establish the first convergence analysis with provable generalization guarantees for SF-DQN with GPI. The theory reveals that SF-DQN with GPI outperforms conventional RL approaches, such as deep Q-network, in terms of both faster convergence rate and better generalization. Numerical experiments on real and synthetic RL tasks support the superior performance of SF-DQN \& GPI, aligning with our theoretical findings.
title SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.15920