When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Ji, Tang, Bin, Xiao, Jian, Cui, Qimei, Li, Xingwang, Quek, Tony Q. S.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913993075982336
author Wang, Ji
Tang, Bin
Xiao, Jian
Cui, Qimei
Li, Xingwang
Quek, Tony Q. S.
author_facet Wang, Ji
Tang, Bin
Xiao, Jian
Cui, Qimei
Li, Xingwang
Quek, Tony Q. S.
contents As the real propagation environment becomes in creasingly complex and dynamic, millimeter wave beam prediction faces huge challenges. However, the powerful cross modal representation capability of vision-language model (VLM) provides a promising approach. The traditional methods that rely on real-time channel state information (CSI) are computationally expensive and often fail to maintain accuracy in such environments. In this paper, we present a VLM-driven contrastive learning based multimodal beam prediction framework that integrates multimodal data via modality-specific encoders. To enforce cross-modal consistency, we adopt a contrastive pretraining strategy to align image and LiDAR features in the latent space. We use location information as text prompts and connect it to the text encoder to introduce language modality, which further improves cross-modal consistency. Experiments on the DeepSense-6G dataset show that our VLM backbone provides additional semantic grounding. Compared with existing methods, the overall distance-based accuracy score (DBA-Score) of 0.9016, corresponding to 1.46% average improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework
Wang, Ji
Tang, Bin
Xiao, Jian
Cui, Qimei
Li, Xingwang
Quek, Tony Q. S.
Signal Processing
As the real propagation environment becomes in creasingly complex and dynamic, millimeter wave beam prediction faces huge challenges. However, the powerful cross modal representation capability of vision-language model (VLM) provides a promising approach. The traditional methods that rely on real-time channel state information (CSI) are computationally expensive and often fail to maintain accuracy in such environments. In this paper, we present a VLM-driven contrastive learning based multimodal beam prediction framework that integrates multimodal data via modality-specific encoders. To enforce cross-modal consistency, we adopt a contrastive pretraining strategy to align image and LiDAR features in the latent space. We use location information as text prompts and connect it to the text encoder to introduce language modality, which further improves cross-modal consistency. Experiments on the DeepSense-6G dataset show that our VLM backbone provides additional semantic grounding. Compared with existing methods, the overall distance-based accuracy score (DBA-Score) of 0.9016, corresponding to 1.46% average improvement.
title When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework
topic Signal Processing
url https://arxiv.org/abs/2508.00456