How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
Fuente:
arXiv
Saved in:
| Main Authors: | Kashyap, Gautam Siddharth, Kulahara, Manaswi, Joshi, Nipun, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification
by: Kulahara, Manaswi, et al.
Published: (2025)
by: Kulahara, Manaswi, et al.
Published: (2025)
Can We Predict Your Next Move Without Breaking Your Privacy?
by: Soni, Arpita, et al.
Published: (2025)
by: Soni, Arpita, et al.
Published: (2025)
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval
by: Dong, Le, et al.
Published: (2024)
by: Dong, Le, et al.
Published: (2024)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025)
by: Amangeldi, Aidar, et al.
Published: (2025)
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
ViT-5: Vision Transformers for The Mid-2020s
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
Sub-token ViT Embedding via Stochastic Resonance Transformers
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
How to train your ViT for OOD Detection
by: Mueller, Maximilian, et al.
Published: (2024)
by: Mueller, Maximilian, et al.
Published: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
by: Ibtehaz, Nabil, et al.
Published: (2024)
by: Ibtehaz, Nabil, et al.
Published: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation
by: Mutlu, Abdulvahap, et al.
Published: (2025)
by: Mutlu, Abdulvahap, et al.
Published: (2025)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification
by: Liu, Hao, et al.
Published: (2025)
by: Liu, Hao, et al.
Published: (2025)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
Deeper Inside Deep ViT
by: Hong, Sungrae
Published: (2025)
by: Hong, Sungrae
Published: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
AResNet-ViT: A Hybrid CNN-Transformer Network for Benign and Malignant Breast Nodule Classification in Ultrasound Images
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
DIOR-ViT: Differential Ordinal Learning Vision Transformer for Cancer Classification in Pathology Images
by: Lee, Ju Cheon, et al.
Published: (2024)
by: Lee, Ju Cheon, et al.
Published: (2024)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
by: Yao, Ting, et al.
Published: (2024)
by: Yao, Ting, et al.
Published: (2024)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Tiny-ViT: A Compact Vision Transformer for Efficient and Explainable Potato Leaf Disease Classification
by: Mia, Shakil, et al.
Published: (2026)
by: Mia, Shakil, et al.
Published: (2026)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
by: Xu, Xuwei, et al.
Published: (2023)
by: Xu, Xuwei, et al.
Published: (2023)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
by: Hernandez, Juan Manuel, et al.
Published: (2026)
by: Hernandez, Juan Manuel, et al.
Published: (2026)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
by: Atzori, Andrea, et al.
Published: (2025)
by: Atzori, Andrea, et al.
Published: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
by: Tai, Yu-Shan, et al.
Published: (2024)
by: Tai, Yu-Shan, et al.
Published: (2024)
Frequency-Adaptive Discrete Cosine-ViT-ResNet Architecture for Sparse-Data Vision
by: Kang, Ziyue, et al.
Published: (2025)
by: Kang, Ziyue, et al.
Published: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
by: Gong, Wenyi, et al.
Published: (2025)
by: Gong, Wenyi, et al.
Published: (2025)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
MagicBathyNet: A Multimodal Remote Sensing Dataset for Bathymetry Prediction and Pixel-based Classification in Shallow Waters
by: Agrafiotis, Panagiotis, et al.
Published: (2024)
by: Agrafiotis, Panagiotis, et al.
Published: (2024)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
ViTCAE: ViT-based Class-conditioned Autoencoder
by: Jebraeeli, Vahid, et al.
Published: (2025)
by: Jebraeeli, Vahid, et al.
Published: (2025)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
by: Dey, Sainath, et al.
Published: (2025)
by: Dey, Sainath, et al.
Published: (2025)
Similar Items
-
Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification
by: Kulahara, Manaswi, et al.
Published: (2025) -
Can We Predict Your Next Move Without Breaking Your Privacy?
by: Soni, Arpita, et al.
Published: (2025) -
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval
by: Dong, Le, et al.
Published: (2024) -
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025) -
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
by: Zhang, Jiaqing, et al.
Published: (2024)