Saved in:
Bibliographic Details
Main Authors: Pahari, Soham, Srinivas, M.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.24404
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915700951482368
author Pahari, Soham
Srinivas, M.
author_facet Pahari, Soham
Srinivas, M.
contents Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This limit their effectiveness in spatial tasks such as visual navigation and geo-localization. This work discuss about the potential scope of this field and eventually propose an idea visual reasoning paradigm Geo-Consistent Visual Planning, our introduced framework called Visual Reasoning for Localization, or ViReLoc, which performs planning and localization using only visual representations. The proposed framework learns spatial dependencies and geometric relations that text based reasoning often suffer to understand. By encoding step by step inference in the visual domain and optimizing with reinforcement based objectives, ViReLoc plans routes between two given ground images. The system also integrates contrastive learning and adaptive feature interaction to align cross view perspectives and reduce viewpoint differences. Experiments across diverse navigation and localization scenarios show consistent improvements in spatial reasoning accuracy and cross view retrieval performance. These results establish visual reasoning as a strong complementary approach for navigation and localization, and show that such tasks can be performed without real time global positioning system data, leading to more secure navigation solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24404
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning
Pahari, Soham
Srinivas, M.
Machine Learning
Computer Vision and Pattern Recognition
68T45, 68T40, 68T20, 93C85
I.2.9; I.4.8; I.5.1; C.2.4
Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This limit their effectiveness in spatial tasks such as visual navigation and geo-localization. This work discuss about the potential scope of this field and eventually propose an idea visual reasoning paradigm Geo-Consistent Visual Planning, our introduced framework called Visual Reasoning for Localization, or ViReLoc, which performs planning and localization using only visual representations. The proposed framework learns spatial dependencies and geometric relations that text based reasoning often suffer to understand. By encoding step by step inference in the visual domain and optimizing with reinforcement based objectives, ViReLoc plans routes between two given ground images. The system also integrates contrastive learning and adaptive feature interaction to align cross view perspectives and reduce viewpoint differences. Experiments across diverse navigation and localization scenarios show consistent improvements in spatial reasoning accuracy and cross view retrieval performance. These results establish visual reasoning as a strong complementary approach for navigation and localization, and show that such tasks can be performed without real time global positioning system data, leading to more secure navigation solutions.
title Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning
topic Machine Learning
Computer Vision and Pattern Recognition
68T45, 68T40, 68T20, 93C85
I.2.9; I.4.8; I.5.1; C.2.4
url https://arxiv.org/abs/2512.24404