Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Haoxiang, Zhang, Li, Zhao, Yu, Yang, Zhou, Cao, Jinghan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913965439713280
author Gao, Haoxiang
Zhang, Li
Zhao, Yu
Yang, Zhou
Cao, Jinghan
author_facet Gao, Haoxiang
Zhang, Li
Zhao, Yu
Yang, Zhou
Cao, Jinghan
contents Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient vehicle deployment. In this paper, we propose a knowledge distillation method that transfers knowledge from large-scale vision-language foundation models to efficient vision networks, and we apply it to pedestrian behavior prediction and scene understanding tasks, achieving promising results in generating more diverse and comprehensive semantic attributes. We also utilize multiple pre-trained models and ensemble techniques to boost the model's performance. We further examined the effectiveness of the model after knowledge distillation; the results show significant metric improvements in open-vocabulary perception and trajectory prediction tasks, which can potentially enhance the end-to-end performance of autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
Gao, Haoxiang
Zhang, Li
Zhao, Yu
Yang, Zhou
Cao, Jinghan
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient vehicle deployment. In this paper, we propose a knowledge distillation method that transfers knowledge from large-scale vision-language foundation models to efficient vision networks, and we apply it to pedestrian behavior prediction and scene understanding tasks, achieving promising results in generating more diverse and comprehensive semantic attributes. We also utilize multiple pre-trained models and ensemble techniques to boost the model's performance. We further examined the effectiveness of the model after knowledge distillation; the results show significant metric improvements in open-vocabulary perception and trajectory prediction tasks, which can potentially enhance the end-to-end performance of autonomous driving.
title Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2501.06680