Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Derun, Li, Changye, Wang, Yue, Ren, Jianwei, Wen, Xin, Li, Pengxiang, Xu, Leimeng, Zhan, Kun, Jia, Peng, Lang, Xianpeng, Xu, Ningyi, Zhao, Hang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909807641886720
author Li, Derun
Li, Changye
Wang, Yue
Ren, Jianwei
Wen, Xin
Li, Pengxiang
Xu, Leimeng
Zhan, Kun
Jia, Peng
Lang, Xianpeng
Xu, Ningyi
Zhao, Hang
author_facet Li, Derun
Li, Changye
Wang, Yue
Ren, Jianwei
Wen, Xin
Li, Pengxiang
Xu, Leimeng
Zhan, Kun
Jia, Peng
Lang, Xianpeng
Xu, Ningyi
Zhao, Hang
contents Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible trajectories, they often fail to capture the nuanced variability of personalized driving styles due to dataset biases and distributional shifts. To address this, we introduce TrajHF, a human feedback-driven finetuning framework for generative trajectory models, designed to align motion planning with diverse driving styles. TrajHF incorporates multi-conditional denoiser and reinforcement learning with human feedback to refine multi-modal trajectory generation beyond conventional imitation learning. This enables better alignment with human driving preferences while maintaining safety and feasibility constraints. TrajHF achieves performance comparable to the state-of-the-art on NavSim benchmark. TrajHF sets a new paradigm for personalized and adaptable trajectory generation in autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback
Li, Derun
Li, Changye
Wang, Yue
Ren, Jianwei
Wen, Xin
Li, Pengxiang
Xu, Leimeng
Zhan, Kun
Jia, Peng
Lang, Xianpeng
Xu, Ningyi
Zhao, Hang
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible trajectories, they often fail to capture the nuanced variability of personalized driving styles due to dataset biases and distributional shifts. To address this, we introduce TrajHF, a human feedback-driven finetuning framework for generative trajectory models, designed to align motion planning with diverse driving styles. TrajHF incorporates multi-conditional denoiser and reinforcement learning with human feedback to refine multi-modal trajectory generation beyond conventional imitation learning. This enables better alignment with human driving preferences while maintaining safety and feasibility constraints. TrajHF achieves performance comparable to the state-of-the-art on NavSim benchmark. TrajHF sets a new paradigm for personalized and adaptable trajectory generation in autonomous driving.
title Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.10434