DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Zhiyi, Ma, Enhui, Li, Fang, Lai, Zhiyi, Ho, Kalok, Wu, Zhanqian, Zhou, Lijun, Chen, Long, Sun, Chitian, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Yu, Kaicheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911052508168192
author Hou, Zhiyi
Ma, Enhui
Li, Fang
Lai, Zhiyi
Ho, Kalok
Wu, Zhanqian
Zhou, Lijun
Chen, Long
Sun, Chitian
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Yu, Kaicheng
author_facet Hou, Zhiyi
Ma, Enhui
Li, Fang
Lai, Zhiyi
Ho, Kalok
Wu, Zhanqian
Zhou, Lijun
Chen, Long
Sun, Chitian
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Yu, Kaicheng
contents Autonomous driving has seen significant progress, driven by extensive real-world data. However, in long-tail scenarios, accurately predicting the safety of the ego vehicle's future motion remains a major challenge due to uncertainties in dynamic environments and limitations in data coverage. In this work, we aim to explore whether it is possible to enhance the motion risk prediction capabilities of Vision-Language Models (VLM) by synthesizing high-risk motion data. Specifically, we introduce a Bird's-Eye View (BEV) based motion simulation method to model risks from three aspects: the ego-vehicle, other vehicles, and the environment. This allows us to synthesize plug-and-play, high-risk motion data suitable for VLM training, which we call DriveMRP-10K. Furthermore, we design a VLM-agnostic motion risk estimation framework, named DriveMRP-Agent. This framework incorporates a novel information injection strategy for global context, ego-vehicle perspective, and trajectory projection, enabling VLMs to effectively reason about the spatial relationships between motion waypoints and the environment. Extensive experiments demonstrate that by fine-tuning with DriveMRP-10K, our DriveMRP-Agent framework can significantly improve the motion risk prediction performance of multiple VLM baselines, with the accident recognition accuracy soaring from 27.13% to 88.03%. Moreover, when tested via zero-shot evaluation on an in-house real-world high-risk motion dataset, DriveMRP-Agent achieves a significant performance leap, boosting the accuracy from base_model's 29.42% to 68.50%, which showcases the strong generalization capabilities of our method in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
Hou, Zhiyi
Ma, Enhui
Li, Fang
Lai, Zhiyi
Ho, Kalok
Wu, Zhanqian
Zhou, Lijun
Chen, Long
Sun, Chitian
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Yu, Kaicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
I.4.8; I.2.7; I.2.10
Autonomous driving has seen significant progress, driven by extensive real-world data. However, in long-tail scenarios, accurately predicting the safety of the ego vehicle's future motion remains a major challenge due to uncertainties in dynamic environments and limitations in data coverage. In this work, we aim to explore whether it is possible to enhance the motion risk prediction capabilities of Vision-Language Models (VLM) by synthesizing high-risk motion data. Specifically, we introduce a Bird's-Eye View (BEV) based motion simulation method to model risks from three aspects: the ego-vehicle, other vehicles, and the environment. This allows us to synthesize plug-and-play, high-risk motion data suitable for VLM training, which we call DriveMRP-10K. Furthermore, we design a VLM-agnostic motion risk estimation framework, named DriveMRP-Agent. This framework incorporates a novel information injection strategy for global context, ego-vehicle perspective, and trajectory projection, enabling VLMs to effectively reason about the spatial relationships between motion waypoints and the environment. Extensive experiments demonstrate that by fine-tuning with DriveMRP-10K, our DriveMRP-Agent framework can significantly improve the motion risk prediction performance of multiple VLM baselines, with the accident recognition accuracy soaring from 27.13% to 88.03%. Moreover, when tested via zero-shot evaluation on an in-house real-world high-risk motion dataset, DriveMRP-Agent achieves a significant performance leap, boosting the accuracy from base_model's 29.42% to 68.50%, which showcases the strong generalization capabilities of our method in real-world scenarios.
title DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
I.4.8; I.2.7; I.2.10
url https://arxiv.org/abs/2507.02948