Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Hao, Gao, Zhanning, Chen, Zhili, Ye, Maosheng, Chen, Qifeng, Cao, Tongyi, Qi, Honggang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915555236118528
author Zhou, Hao
Gao, Zhanning
Chen, Zhili
Ye, Maosheng
Chen, Qifeng
Cao, Tongyi
Qi, Honggang
author_facet Zhou, Hao
Gao, Zhanning
Chen, Zhili
Ye, Maosheng
Chen, Qifeng
Cao, Tongyi
Qi, Honggang
contents In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hints of Prompt (HoP) framework, which introduces three key enhancements: Affinity hint to emphasize instance-level structure by strengthening token-wise connections, Semantic hint to incorporate high-level information relevant to driving-specific cases, such as complex interactions among vehicles and traffic signs, and Question hint to align visual features with the query context, focusing on question-relevant regions. These hints are fused through a Hint Fusion module, enriching visual representations by capturing driving-related representations with limited domain data, ensuring faster adaptation to driving scenarios. Extensive experiments confirm the effectiveness of the HoP framework, showing that it significantly outperforms previous state-of-the-art methods in all key metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13076
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
Zhou, Hao
Gao, Zhanning
Chen, Zhili
Ye, Maosheng
Chen, Qifeng
Cao, Tongyi
Qi, Honggang
Computer Vision and Pattern Recognition
In light of the dynamic nature of autonomous driving environments and stringent safety requirements, general MLLMs combined with CLIP alone often struggle to accurately represent driving-specific scenarios, particularly in complex interactions and long-tail cases. To address this, we propose the Hints of Prompt (HoP) framework, which introduces three key enhancements: Affinity hint to emphasize instance-level structure by strengthening token-wise connections, Semantic hint to incorporate high-level information relevant to driving-specific cases, such as complex interactions among vehicles and traffic signs, and Question hint to align visual features with the query context, focusing on question-relevant regions. These hints are fused through a Hint Fusion module, enriching visual representations by capturing driving-related representations with limited domain data, ensuring faster adaptation to driving scenarios. Extensive experiments confirm the effectiveness of the HoP framework, showing that it significantly outperforms previous state-of-the-art methods in all key metrics.
title Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13076