MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914251561500672 |
|---|---|
| author | Chaudhary, Aditya Barman, Sneha Singha, Mainak Jha, Ankit Mishra, Girish Banerjee, Biplab |
| author_facet | Chaudhary, Aditya Barman, Sneha Singha, Mainak Jha, Ankit Mishra, Girish Banerjee, Biplab |
| contents | In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_08420 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP Chaudhary, Aditya Barman, Sneha Singha, Mainak Jha, Ankit Mishra, Girish Banerjee, Biplab Computer Vision and Pattern Recognition In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI. |
| title | MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.08420 |