Saved in:
Bibliographic Details
Main Authors: Jia, Furong, Dai, Ling, Deng, Wenjin, Zhang, Fan, Hu, Chen, Jiang, Daxin, Liu, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.09463
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908859933655040
author Jia, Furong
Dai, Ling
Deng, Wenjin
Zhang, Fan
Hu, Chen
Jiang, Daxin
Liu, Yu
author_facet Jia, Furong
Dai, Ling
Deng, Wenjin
Zhang, Fan
Hu, Chen
Jiang, Daxin
Liu, Yu
contents Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are sparse, long-tailed, and highly ambiguous. Previous approaches, bound by internal knowledge, often fail to provide verifiable results, yielding confident but ungrounded predictions when faced with confounded evidence. To address these challenges, we propose SpotAgent, a framework that formalizes geo-localization into an agentic reasoning process that leverages expert-level reasoning to synergize visual interpretation with tool-assisted verification. SpotAgent actively explores and verifies visual cues by leveraging external tools (e.g., web search, maps) through a ReAct diagram. We introduce a 3-stage post-training pipeline starting with a Supervised Fine-Tuning (SFT) stage for basic alignment, followed by an Agentic Cold Start phase utilizing high-quality trajectories synthesized via a Multi-Agent framework, aiming to instill tool-calling expertise. Subsequently, the model's reasoning capabilities are refined through Reinforcement Learning. We propose a Spatially-Aware Dynamic Filtering strategy to enhance the efficiency of the RL stage by prioritizing learnable samples based on spatial difficulty. Extensive experiments on standard benchmarks demonstrate that SpotAgent achieves state-of-the-art performance, effectively mitigating hallucinations while delivering precise and verifiable geo-localization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09463
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
Jia, Furong
Dai, Ling
Deng, Wenjin
Zhang, Fan
Hu, Chen
Jiang, Daxin
Liu, Yu
Artificial Intelligence
Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are sparse, long-tailed, and highly ambiguous. Previous approaches, bound by internal knowledge, often fail to provide verifiable results, yielding confident but ungrounded predictions when faced with confounded evidence. To address these challenges, we propose SpotAgent, a framework that formalizes geo-localization into an agentic reasoning process that leverages expert-level reasoning to synergize visual interpretation with tool-assisted verification. SpotAgent actively explores and verifies visual cues by leveraging external tools (e.g., web search, maps) through a ReAct diagram. We introduce a 3-stage post-training pipeline starting with a Supervised Fine-Tuning (SFT) stage for basic alignment, followed by an Agentic Cold Start phase utilizing high-quality trajectories synthesized via a Multi-Agent framework, aiming to instill tool-calling expertise. Subsequently, the model's reasoning capabilities are refined through Reinforcement Learning. We propose a Spatially-Aware Dynamic Filtering strategy to enhance the efficiency of the RL stage by prioritizing learnable samples based on spatial difficulty. Extensive experiments on standard benchmarks demonstrate that SpotAgent achieves state-of-the-art performance, effectively mitigating hallucinations while delivering precise and verifiable geo-localization.
title SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2602.09463