GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yikun, Liu, Zuyan, Wang, Ziyi, Hu, Han, Liu, Pengfei, Rao, Yongming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917153034207232
author Wang, Yikun
Liu, Zuyan
Wang, Ziyi
Hu, Han
Liu, Pengfei
Rao, Yongming
author_facet Wang, Yikun
Liu, Zuyan
Wang, Ziyi
Hu, Han
Liu, Pengfei
Rao, Yongming
contents Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization task, which requires not only nuanced visual grounding but also web search to confirm or refine hypotheses during reasoning. Since existing geolocalization benchmarks fail to meet the need for high-resolution imagery and the localization challenge for deep agentic reasoning, we curate GeoBench, a benchmark that includes photos and panoramas from around the world, along with a subset of satellite images of different cities to rigorously evaluate the geolocalization ability of agentic models. We also propose GeoVista, an agentic model that seamlessly integrates tool invocation within the reasoning loop, including an image-zoom-in tool to magnify regions of interest and a web-search tool to retrieve related web information. We develop a complete training pipeline for it, including a cold-start supervised fine-tuning (SFT) stage to learn reasoning patterns and tool-use priors, followed by a reinforcement learning (RL) stage to further enhance reasoning ability. We adopt a hierarchical reward to leverage multi-level geographical information and improve overall geolocalization performance. Experimental results show that GeoVista surpasses other open-source agentic models on the geolocalization task greatly and achieves performance comparable to closed-source models such as Gemini-2.5-flash and GPT-5 on most metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15705
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
Wang, Yikun
Liu, Zuyan
Wang, Ziyi
Hu, Han
Liu, Pengfei
Rao, Yongming
Computer Vision and Pattern Recognition
Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization task, which requires not only nuanced visual grounding but also web search to confirm or refine hypotheses during reasoning. Since existing geolocalization benchmarks fail to meet the need for high-resolution imagery and the localization challenge for deep agentic reasoning, we curate GeoBench, a benchmark that includes photos and panoramas from around the world, along with a subset of satellite images of different cities to rigorously evaluate the geolocalization ability of agentic models. We also propose GeoVista, an agentic model that seamlessly integrates tool invocation within the reasoning loop, including an image-zoom-in tool to magnify regions of interest and a web-search tool to retrieve related web information. We develop a complete training pipeline for it, including a cold-start supervised fine-tuning (SFT) stage to learn reasoning patterns and tool-use priors, followed by a reinforcement learning (RL) stage to further enhance reasoning ability. We adopt a hierarchical reward to leverage multi-level geographical information and improve overall geolocalization performance. Experimental results show that GeoVista surpasses other open-source agentic models on the geolocalization task greatly and achieves performance comparable to closed-source models such as Gemini-2.5-flash and GPT-5 on most metrics.
title GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15705