AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Shixiong, Zhang, Chenghao, Fan, Lubin, Zhou, Yuan, Fan, Bin, Xiang, Shiming, Meng, Gaofeng, Ye, Jieping
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909736942698496
author Xu, Shixiong
Zhang, Chenghao
Fan, Lubin
Zhou, Yuan
Fan, Bin
Xiang, Shiming
Meng, Gaofeng
Ye, Jieping
author_facet Xu, Shixiong
Zhang, Chenghao
Fan, Lubin
Zhou, Yuan
Fan, Bin
Xiang, Shiming
Meng, Gaofeng
Ye, Jieping
contents Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Xu, Shixiong
Zhang, Chenghao
Fan, Lubin
Zhou, Yuan
Fan, Bin
Xiang, Shiming
Meng, Gaofeng
Ye, Jieping
Computer Vision and Pattern Recognition
Artificial Intelligence
Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively.
title AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.10667