Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Biao, Fang, Meng, Chen, Ling, Xu, Ke, Cheng, Tao, Wang, Jun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908748205785088
author Wu, Biao
Fang, Meng
Chen, Ling
Xu, Ke
Cheng, Tao
Wang, Jun
author_facet Wu, Biao
Fang, Meng
Chen, Ling
Xu, Ke
Cheng, Tao
Wang, Jun
contents Recent advances in vision-language models have opened up new possibilities for reasoning-driven image geolocalization. However, existing approaches often rely on synthetic reasoning annotations or external image retrieval, which can limit interpretability and generalizability. In this paper, we present Geo-R, a retrieval-free framework that uncovers structured reasoning paths from existing ground-truth coordinates and optimizes geolocation accuracy via reinforcement learning. We propose the Chain of Region, a rule-based hierarchical reasoning paradigm that generates precise, interpretable supervision by mapping GPS coordinates to geographic entities (e.g., country, province, city) without relying on model-generated or synthetic labels. Building on this, we introduce a lightweight reinforcement learning strategy with coordinate-aligned rewards based on Haversine distance, enabling the model to refine predictions through spatially meaningful feedback. Our approach bridges structured geographic reasoning with direct spatial supervision, yielding improved localization accuracy, stronger generalization, and more transparent inference. Experimental results across multiple benchmarks confirm the effectiveness of Geo-R, establishing a new retrieval-free paradigm for scalable and interpretable image geolocalization. To facilitate further research and ensure reproducibility, both the model and code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00388
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
Wu, Biao
Fang, Meng
Chen, Ling
Xu, Ke
Cheng, Tao
Wang, Jun
Computation and Language
F.2.2; I.2.7
Recent advances in vision-language models have opened up new possibilities for reasoning-driven image geolocalization. However, existing approaches often rely on synthetic reasoning annotations or external image retrieval, which can limit interpretability and generalizability. In this paper, we present Geo-R, a retrieval-free framework that uncovers structured reasoning paths from existing ground-truth coordinates and optimizes geolocation accuracy via reinforcement learning. We propose the Chain of Region, a rule-based hierarchical reasoning paradigm that generates precise, interpretable supervision by mapping GPS coordinates to geographic entities (e.g., country, province, city) without relying on model-generated or synthetic labels. Building on this, we introduce a lightweight reinforcement learning strategy with coordinate-aligned rewards based on Haversine distance, enabling the model to refine predictions through spatially meaningful feedback. Our approach bridges structured geographic reasoning with direct spatial supervision, yielding improved localization accuracy, stronger generalization, and more transparent inference. Experimental results across multiple benchmarks confirm the effectiveness of Geo-R, establishing a new retrieval-free paradigm for scalable and interpretable image geolocalization. To facilitate further research and ensure reproducibility, both the model and code will be made publicly available.
title Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
topic Computation and Language
F.2.2; I.2.7
url https://arxiv.org/abs/2601.00388