How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kashyap, Gautam Siddharth, Kulahara, Manaswi, Joshi, Nipun, Naseem, Usman
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917158796132352
author Kashyap, Gautam Siddharth
Kulahara, Manaswi
Joshi, Nipun
Naseem, Usman
author_facet Kashyap, Gautam Siddharth
Kulahara, Manaswi
Joshi, Nipun
Naseem, Usman
contents Remote sensing datasets offer significant promise for tackling key classification tasks such as land-use categorization, object presence detection, and rural/urban classification. However, many existing studies tend to focus on narrow tasks or datasets, which limits their ability to generalize across various remote sensing classification challenges. To overcome this, we propose a novel model, SpatialNet-ViT, leveraging the power of Vision Transformers (ViTs) and Multi-Task Learning (MTL). This integrated approach combines spatial awareness with contextual understanding, improving both classification accuracy and scalability. Additionally, techniques like data augmentation, transfer learning, and multi-task learning are employed to enhance model robustness and its ability to generalize across diverse datasets
format Preprint
id arxiv_https___arxiv_org_abs_2506_22501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
Kashyap, Gautam Siddharth
Kulahara, Manaswi
Joshi, Nipun
Naseem, Usman
Computer Vision and Pattern Recognition
Artificial Intelligence
Remote sensing datasets offer significant promise for tackling key classification tasks such as land-use categorization, object presence detection, and rural/urban classification. However, many existing studies tend to focus on narrow tasks or datasets, which limits their ability to generalize across various remote sensing classification challenges. To overcome this, we propose a novel model, SpatialNet-ViT, leveraging the power of Vision Transformers (ViTs) and Multi-Task Learning (MTL). This integrated approach combines spatial awareness with contextual understanding, improving both classification accuracy and scalability. Additionally, techniques like data augmentation, transfer learning, and multi-task learning are employed to enhance model robustness and its ability to generalize across diverse datasets
title How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.22501