How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917158796132352 |
|---|---|
| author | Kashyap, Gautam Siddharth Kulahara, Manaswi Joshi, Nipun Naseem, Usman |
| author_facet | Kashyap, Gautam Siddharth Kulahara, Manaswi Joshi, Nipun Naseem, Usman |
| contents | Remote sensing datasets offer significant promise for tackling key classification tasks such as land-use categorization, object presence detection, and rural/urban classification. However, many existing studies tend to focus on narrow tasks or datasets, which limits their ability to generalize across various remote sensing classification challenges. To overcome this, we propose a novel model, SpatialNet-ViT, leveraging the power of Vision Transformers (ViTs) and Multi-Task Learning (MTL). This integrated approach combines spatial awareness with contextual understanding, improving both classification accuracy and scalability. Additionally, techniques like data augmentation, transfer learning, and multi-task learning are employed to enhance model robustness and its ability to generalize across diverse datasets |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_22501 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT? Kashyap, Gautam Siddharth Kulahara, Manaswi Joshi, Nipun Naseem, Usman Computer Vision and Pattern Recognition Artificial Intelligence Remote sensing datasets offer significant promise for tackling key classification tasks such as land-use categorization, object presence detection, and rural/urban classification. However, many existing studies tend to focus on narrow tasks or datasets, which limits their ability to generalize across various remote sensing classification challenges. To overcome this, we propose a novel model, SpatialNet-ViT, leveraging the power of Vision Transformers (ViTs) and Multi-Task Learning (MTL). This integrated approach combines spatial awareness with contextual understanding, improving both classification accuracy and scalability. Additionally, techniques like data augmentation, transfer learning, and multi-task learning are employed to enhance model robustness and its ability to generalize across diverse datasets |
| title | How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT? |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2506.22501 |