Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Hanlin, Yao, Jian, Liu, Weiming, Wang, Qing, Qin, Hanmin, Kong, Hansheng, Tang, Kirk, Xiong, Jiechao, Yu, Chao, Li, Kai, Xing, Junliang, Chen, Hongwu, Zhuo, Juchao, Fu, Qiang, Wei, Yang, Fu, Haobo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916448573587456
author Yang, Hanlin
Yao, Jian
Liu, Weiming
Wang, Qing
Qin, Hanmin
Kong, Hansheng
Tang, Kirk
Xiong, Jiechao
Yu, Chao
Li, Kai
Xing, Junliang
Chen, Hongwu
Zhuo, Juchao
Fu, Qiang
Wei, Yang
Fu, Haobo
author_facet Yang, Hanlin
Yao, Jian
Liu, Weiming
Wang, Qing
Qin, Hanmin
Kong, Hansheng
Tang, Kirk
Xiong, Jiechao
Yu, Chao
Li, Kai
Xing, Junliang
Chen, Hongwu
Zhuo, Juchao
Fu, Qiang
Wei, Yang
Fu, Haobo
contents Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse policies recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovery. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information. This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style. We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse policies from expert data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15910
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
Yang, Hanlin
Yao, Jian
Liu, Weiming
Wang, Qing
Qin, Hanmin
Kong, Hansheng
Tang, Kirk
Xiong, Jiechao
Yu, Chao
Li, Kai
Xing, Junliang
Chen, Hongwu
Zhuo, Juchao
Fu, Qiang
Wei, Yang
Fu, Haobo
Machine Learning
Artificial Intelligence
Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse policies recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovery. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information. This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style. We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse policies from expert data.
title Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.15910