Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909807641886720 |
|---|---|
| author | Li, Derun Li, Changye Wang, Yue Ren, Jianwei Wen, Xin Li, Pengxiang Xu, Leimeng Zhan, Kun Jia, Peng Lang, Xianpeng Xu, Ningyi Zhao, Hang |
| author_facet | Li, Derun Li, Changye Wang, Yue Ren, Jianwei Wen, Xin Li, Pengxiang Xu, Leimeng Zhan, Kun Jia, Peng Lang, Xianpeng Xu, Ningyi Zhao, Hang |
| contents | Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible trajectories, they often fail to capture the nuanced variability of personalized driving styles due to dataset biases and distributional shifts. To address this, we introduce TrajHF, a human feedback-driven finetuning framework for generative trajectory models, designed to align motion planning with diverse driving styles. TrajHF incorporates multi-conditional denoiser and reinforcement learning with human feedback to refine multi-modal trajectory generation beyond conventional imitation learning. This enables better alignment with human driving preferences while maintaining safety and feasibility constraints. TrajHF achieves performance comparable to the state-of-the-art on NavSim benchmark. TrajHF sets a new paradigm for personalized and adaptable trajectory generation in autonomous driving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10434 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback Li, Derun Li, Changye Wang, Yue Ren, Jianwei Wen, Xin Li, Pengxiang Xu, Leimeng Zhan, Kun Jia, Peng Lang, Xianpeng Xu, Ningyi Zhao, Hang Robotics Computer Vision and Pattern Recognition Machine Learning Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible trajectories, they often fail to capture the nuanced variability of personalized driving styles due to dataset biases and distributional shifts. To address this, we introduce TrajHF, a human feedback-driven finetuning framework for generative trajectory models, designed to align motion planning with diverse driving styles. TrajHF incorporates multi-conditional denoiser and reinforcement learning with human feedback to refine multi-modal trajectory generation beyond conventional imitation learning. This enables better alignment with human driving preferences while maintaining safety and feasibility constraints. TrajHF achieves performance comparable to the state-of-the-art on NavSim benchmark. TrajHF sets a new paradigm for personalized and adaptable trajectory generation in autonomous driving. |
| title | Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback |
| topic | Robotics Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2503.10434 |