AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909736942698496 |
|---|---|
| author | Xu, Shixiong Zhang, Chenghao Fan, Lubin Zhou, Yuan Fan, Bin Xiang, Shiming Meng, Gaofeng Ye, Jieping |
| author_facet | Xu, Shixiong Zhang, Chenghao Fan, Lubin Zhou, Yuan Fan, Bin Xiang, Shiming Meng, Gaofeng Ye, Jieping |
| contents | Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_10667 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models Xu, Shixiong Zhang, Chenghao Fan, Lubin Zhou, Yuan Fan, Bin Xiang, Shiming Meng, Gaofeng Ye, Jieping Computer Vision and Pattern Recognition Artificial Intelligence Large visual language models (LVLMs) have demonstrated impressive performance in coarse-grained geo-localization at the country or city level, but they struggle with fine-grained street-level localization within urban areas. In this paper, we explore integrating city-wide address localization capabilities into LVLMs, facilitating flexible address-related question answering using street-view images. A key challenge is that the street-view visual question-and-answer (VQA) data provides only microscopic visual cues, leading to subpar performance in fine-tuned models. To tackle this issue, we incorporate perspective-invariant satellite images as macro cues and propose cross-view alignment tuning including a satellite-view and street-view image grafting mechanism, along with an automatic label generation mechanism. Then LVLM's global understanding of street distribution is enhanced through cross-view matching. Our proposed model, named AddressVLM, consists of two-stage training protocols: cross-view alignment tuning and address localization tuning. Furthermore, we have constructed two street-view VQA datasets based on image address localization datasets from Pittsburgh and San Francisco. Qualitative and quantitative evaluations demonstrate that AddressVLM outperforms counterpart LVLMs by over 9% and 12% in average address localization accuracy on these two datasets, respectively. |
| title | AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2508.10667 |