GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuhao, Al-Kindi, Sadeer, Veeraraghavan, Ashok, Balakrishnan, Guha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When are Diffusion Priors Helpful in Sparse Reconstruction? A Study with Sparse-view CT
by: Cheung, Matt Y., et al.
Published: (2025)
by: Cheung, Matt Y., et al.
Published: (2025)
COMPASS: Robust Feature Conformal Prediction for Medical Segmentation Metrics
by: Cheung, Matt Y., et al.
Published: (2025)
by: Cheung, Matt Y., et al.
Published: (2025)
Efficient Conformal Volumetry for Template-Based Segmentation
by: Cheung, Matt Y., et al.
Published: (2026)
by: Cheung, Matt Y., et al.
Published: (2026)
Post-Hurricane Debris Segmentation Using Fine-Tuned Foundational Vision Models
by: Amini, Kooshan, et al.
Published: (2025)
by: Amini, Kooshan, et al.
Published: (2025)
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
by: Vyas, Kushal, et al.
Published: (2025)
by: Vyas, Kushal, et al.
Published: (2025)
Metric-Guided Conformal Bounds for Probabilistic Image Reconstruction
by: Cheung, Matt Y, et al.
Published: (2024)
by: Cheung, Matt Y, et al.
Published: (2024)
The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representations
by: Vyas, Kushal, et al.
Published: (2026)
by: Vyas, Kushal, et al.
Published: (2026)
NeRT: Implicit Neural Representations for General Unsupervised Turbulence Mitigation
by: Jiang, Weiyun, et al.
Published: (2023)
by: Jiang, Weiyun, et al.
Published: (2023)
WriteViT: Handwritten Text Generation with Vision Transformer
by: Nam, Dang Hoai, et al.
Published: (2025)
by: Nam, Dang Hoai, et al.
Published: (2025)
Resilient Vision-Tabular Multimodal Learning under Modality Missingness
by: Caruso, Camillo Maria, et al.
Published: (2026)
by: Caruso, Camillo Maria, et al.
Published: (2026)
TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
by: Siampou, Maria Despoina, et al.
Published: (2026)
by: Siampou, Maria Despoina, et al.
Published: (2026)
Learning Transferable Features for Implicit Neural Representations
by: Vyas, Kushal, et al.
Published: (2024)
by: Vyas, Kushal, et al.
Published: (2024)
ViTGAN: Training GANs with Vision Transformers
by: Lee, Kwonjoon, et al.
Published: (2021)
by: Lee, Kwonjoon, et al.
Published: (2021)
SplineCam: Exact Visualization and Characterization of Deep Network Geometry and Decision Boundaries
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2023)
Early Prediction of Type 2 Diabetes Using Multimodal data and Tabular Transformers
by: Khan, Sulaiman, et al.
Published: (2026)
by: Khan, Sulaiman, et al.
Published: (2026)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
by: Mantes, Albert Dominguez, et al.
Published: (2026)
by: Mantes, Albert Dominguez, et al.
Published: (2026)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
ViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers
by: Ozgur, Guray, et al.
Published: (2026)
by: Ozgur, Guray, et al.
Published: (2026)
Fast Amortized Fitting of Scientific Signals Across Time and Ensembles via Transferable Neural Fields
by: Zorek, Sophia, et al.
Published: (2026)
by: Zorek, Sophia, et al.
Published: (2026)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
ViT-MUL: A Baseline Study on Recent Machine Unlearning Methods Applied to Vision Transformers
by: Cho, Ikhyun, et al.
Published: (2024)
by: Cho, Ikhyun, et al.
Published: (2024)
LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
Federated EndoViT: Pretraining Vision Transformers via Federated Learning on Endoscopic Image Collections
by: Kirchner, Max, et al.
Published: (2025)
by: Kirchner, Max, et al.
Published: (2025)
BornoViT: A Novel Efficient Vision Transformer for Bengali Handwritten Basic Characters Classification
by: Chowdhury, Rafi Hassan, et al.
Published: (2026)
by: Chowdhury, Rafi Hassan, et al.
Published: (2026)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
ScriptViT: Vision Transformer-Based Personalized Handwriting Generation
by: Acharya, Sajjan, et al.
Published: (2025)
by: Acharya, Sajjan, et al.
Published: (2025)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
GViT: Representing Images as Gaussians for Visual Recognition
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
VisTabNet: Adapting Vision Transformers for Tabular Data
by: Wydmański, Witold, et al.
Published: (2024)
by: Wydmański, Witold, et al.
Published: (2024)
Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
A Manifold Representation of the Key in Vision Transformers
by: Meng, Li, et al.
Published: (2024)
by: Meng, Li, et al.
Published: (2024)
ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers
by: Lygerakis, Fotios, et al.
Published: (2025)
by: Lygerakis, Fotios, et al.
Published: (2025)
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
by: Ataiefard, Foozhan, et al.
Published: (2024)
by: Ataiefard, Foozhan, et al.
Published: (2024)
TRecViT: A Recurrent Video Transformer
by: Pătrăucean, Viorica, et al.
Published: (2024)
by: Pătrăucean, Viorica, et al.
Published: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
ZACH-ViT: A Zero-Token Vision Transformer with ShuffleStrides Data Augmentation for Robust Lung Ultrasound Classification
by: Angelakis, Athanasios, et al.
Published: (2025)
by: Angelakis, Athanasios, et al.
Published: (2025)
S-E Pipeline: A Vision Transformer (ViT) based Resilient Classification Pipeline for Medical Imaging Against Adversarial Attacks
by: S, Neha A, et al.
Published: (2024)
by: S, Neha A, et al.
Published: (2024)
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning
by: Raju, S M Taslim Uddin, et al.
Published: (2025)
by: Raju, S M Taslim Uddin, et al.
Published: (2025)
ZAYAN: Disentangled Contrastive Transformer for Tabular Remote Sensing Data
by: Habib, Al Zadid Sultan Bin, et al.
Published: (2026)
by: Habib, Al Zadid Sultan Bin, et al.
Published: (2026)
Similar Items
-
When are Diffusion Priors Helpful in Sparse Reconstruction? A Study with Sparse-view CT
by: Cheung, Matt Y., et al.
Published: (2025) -
COMPASS: Robust Feature Conformal Prediction for Medical Segmentation Metrics
by: Cheung, Matt Y., et al.
Published: (2025) -
Efficient Conformal Volumetry for Template-Based Segmentation
by: Cheung, Matt Y., et al.
Published: (2026) -
Post-Hurricane Debris Segmentation Using Fine-Tuned Foundational Vision Models
by: Amini, Kooshan, et al.
Published: (2025) -
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
by: Vyas, Kushal, et al.
Published: (2025)