MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhary, Aditya, Barman, Sneha, Singha, Mainak, Jha, Ankit, Mishra, Girish, Banerjee, Biplab
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914251561500672
author Chaudhary, Aditya
Barman, Sneha
Singha, Mainak
Jha, Ankit
Mishra, Girish
Banerjee, Biplab
author_facet Chaudhary, Aditya
Barman, Sneha
Singha, Mainak
Jha, Ankit
Mishra, Girish
Banerjee, Biplab
contents In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08420
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
Chaudhary, Aditya
Barman, Sneha
Singha, Mainak
Jha, Ankit
Mishra, Girish
Banerjee, Biplab
Computer Vision and Pattern Recognition
In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.
title MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.08420