Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jian-Ting, Chen, Yu-Cheng, Hsieh, Ping-Chun, Ho, Kuo-Hao, Huang, Po-Wei, Wu, Ti-Rong, Wu, I-Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909913380290560
author Guo, Jian-Ting
Chen, Yu-Cheng
Hsieh, Ping-Chun
Ho, Kuo-Hao
Huang, Po-Wei
Wu, Ti-Rong
Wu, I-Chen
author_facet Guo, Jian-Ting
Chen, Yu-Cheng
Hsieh, Ping-Chun
Ho, Kuo-Hao
Huang, Po-Wei
Wu, Ti-Rong
Wu, I-Chen
contents Human-like agents have long been one of the goals in pursuing artificial intelligence. Although reinforcement learning (RL) has achieved superhuman performance in many domains, relatively little attention has been focused on designing human-like RL agents. As a result, many reward-driven RL agents often exhibit unnatural behaviors compared to humans, raising concerns for both interpretability and trustworthiness. To achieve human-like behavior in RL, this paper first formulates human-likeness as trajectory optimization, where the objective is to find an action sequence that closely aligns with human behavior while also maximizing rewards, and adapts the classic receding-horizon control to human-like learning as a tractable and efficient implementation. To achieve this, we introduce Macro Action Quantization (MAQ), a human-like RL framework that distills human demonstrations into macro actions via Vector-Quantized VAE. Experiments on D4RL Adroit benchmarks show that MAQ significantly improves human-likeness, increasing trajectory similarity scores, and achieving the highest human-likeness rankings among all RL agents in the human evaluation study. Our results also demonstrate that MAQ can be easily integrated into various off-the-shelf RL algorithms, opening a promising direction for learning human-like RL agents. Our code is available at https://rlg.iis.sinica.edu.tw/papers/MAQ.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
Guo, Jian-Ting
Chen, Yu-Cheng
Hsieh, Ping-Chun
Ho, Kuo-Hao
Huang, Po-Wei
Wu, Ti-Rong
Wu, I-Chen
Artificial Intelligence
Machine Learning
Robotics
Human-like agents have long been one of the goals in pursuing artificial intelligence. Although reinforcement learning (RL) has achieved superhuman performance in many domains, relatively little attention has been focused on designing human-like RL agents. As a result, many reward-driven RL agents often exhibit unnatural behaviors compared to humans, raising concerns for both interpretability and trustworthiness. To achieve human-like behavior in RL, this paper first formulates human-likeness as trajectory optimization, where the objective is to find an action sequence that closely aligns with human behavior while also maximizing rewards, and adapts the classic receding-horizon control to human-like learning as a tractable and efficient implementation. To achieve this, we introduce Macro Action Quantization (MAQ), a human-like RL framework that distills human demonstrations into macro actions via Vector-Quantized VAE. Experiments on D4RL Adroit benchmarks show that MAQ significantly improves human-likeness, increasing trajectory similarity scores, and achieving the highest human-likeness rankings among all RL agents in the human evaluation study. Our results also demonstrate that MAQ can be easily integrated into various off-the-shelf RL algorithms, opening a promising direction for learning human-like RL agents. Our code is available at https://rlg.iis.sinica.edu.tw/papers/MAQ.
title Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2511.15055