Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913965439713280 |
|---|---|
| author | Gao, Haoxiang Zhang, Li Zhao, Yu Yang, Zhou Cao, Jinghan |
| author_facet | Gao, Haoxiang Zhang, Li Zhao, Yu Yang, Zhou Cao, Jinghan |
| contents | Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient vehicle deployment. In this paper, we propose a knowledge distillation method that transfers knowledge from large-scale vision-language foundation models to efficient vision networks, and we apply it to pedestrian behavior prediction and scene understanding tasks, achieving promising results in generating more diverse and comprehensive semantic attributes. We also utilize multiple pre-trained models and ensemble techniques to boost the model's performance. We further examined the effectiveness of the model after knowledge distillation; the results show significant metric improvements in open-vocabulary perception and trajectory prediction tasks, which can potentially enhance the end-to-end performance of autonomous driving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_06680 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving Gao, Haoxiang Zhang, Li Zhao, Yu Yang, Zhou Cao, Jinghan Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient vehicle deployment. In this paper, we propose a knowledge distillation method that transfers knowledge from large-scale vision-language foundation models to efficient vision networks, and we apply it to pedestrian behavior prediction and scene understanding tasks, achieving promising results in generating more diverse and comprehensive semantic attributes. We also utilize multiple pre-trained models and ensemble techniques to boost the model's performance. We further examined the effectiveness of the model after knowledge distillation; the results show significant metric improvements in open-vocabulary perception and trajectory prediction tasks, which can potentially enhance the end-to-end performance of autonomous driving. |
| title | Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics |
| url | https://arxiv.org/abs/2501.06680 |