GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hacheme, Gilles Quentin, Tadesse, Girmaw Abebe, Robinson, Caleb, Zaytar, Akram, Dodhia, Rahul, Ferres, Juan M. Lavista
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916767725518848
author Hacheme, Gilles Quentin
Tadesse, Girmaw Abebe
Robinson, Caleb
Zaytar, Akram
Dodhia, Rahul
Ferres, Juan M. Lavista
author_facet Hacheme, Gilles Quentin
Tadesse, Girmaw Abebe
Robinson, Caleb
Zaytar, Akram
Dodhia, Rahul
Ferres, Juan M. Lavista
contents Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS-CLIP) that claim zero-shot classification capabilities for satellite imagery nonetheless rely on task-specific pretraining and adaptation to reach competitive performance. We introduce GeoVision Labeler (GVL), a strictly zero-shot classification framework: a vision Large Language Model (vLLM) generates rich, human-readable image descriptions, which are then mapped to user-defined classes by a conventional Large Language Model (LLM). This modular, and interpretable pipeline enables flexible image classification for a large range of use cases. We evaluated GVL across three benchmarks-SpaceNet v7, UC Merced, and RESISC45. It achieves up to 93.2% zero-shot accuracy on the binary Buildings vs. No Buildings task on SpaceNet v7. For complex multi-class classification tasks (UC Merced, RESISC45), we implemented a recursive LLM-driven clustering to form meta-classes at successive depths, followed by hierarchical classification-first resolving coarse groups, then finer distinctions-to deliver competitive zero-shot performance. GVL is open-sourced at https://github.com/microsoft/geo-vision-labeler to catalyze adoption in real-world geospatial workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
Hacheme, Gilles Quentin
Tadesse, Girmaw Abebe
Robinson, Caleb
Zaytar, Akram
Dodhia, Rahul
Ferres, Juan M. Lavista
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
I.2.10; I.2.7; I.4.8; I.5.3
Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS-CLIP) that claim zero-shot classification capabilities for satellite imagery nonetheless rely on task-specific pretraining and adaptation to reach competitive performance. We introduce GeoVision Labeler (GVL), a strictly zero-shot classification framework: a vision Large Language Model (vLLM) generates rich, human-readable image descriptions, which are then mapped to user-defined classes by a conventional Large Language Model (LLM). This modular, and interpretable pipeline enables flexible image classification for a large range of use cases. We evaluated GVL across three benchmarks-SpaceNet v7, UC Merced, and RESISC45. It achieves up to 93.2% zero-shot accuracy on the binary Buildings vs. No Buildings task on SpaceNet v7. For complex multi-class classification tasks (UC Merced, RESISC45), we implemented a recursive LLM-driven clustering to form meta-classes at successive depths, followed by hierarchical classification-first resolving coarse groups, then finer distinctions-to deliver competitive zero-shot performance. GVL is open-sourced at https://github.com/microsoft/geo-vision-labeler to catalyze adoption in real-world geospatial workflows.
title GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
I.2.10; I.2.7; I.4.8; I.5.3
url https://arxiv.org/abs/2505.24340