Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Fengyi, Ma, Jun, Qiu, Waishan, Guo, Cui, Cheng, Jack C. P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911602364645376
author Xu, Fengyi
Ma, Jun
Qiu, Waishan
Guo, Cui
Cheng, Jack C. P.
author_facet Xu, Fengyi
Ma, Jun
Qiu, Waishan
Guo, Cui
Cheng, Jack C. P.
contents Crowdsourced social media imagery provides real-time visual evidence of urban flooding but often lacks reliable geographic metadata for emergency response. Existing Visual Place Recognition (VPR) models struggle to geo-localize these images due to cross-source domain shifts and visual distortions. We present VPR-AttLLM, a model-agnostic framework integrating the semantic reasoning and geospatial knowledge of Large Language Models (LLMs) into VPR pipelines via attention-guided descriptor enhancement. VPR-AttLLM uses LLMs to isolate location-informative regions and suppress transient noise, improving retrieval without model retraining or new data. We evaluate this framework across San Francisco and Hong Kong using established queries, synthetic flooding scenarios, and real social media flood images. Integrating VPR-AttLLM with state-of-the-art models (CosPlace, EigenPlaces, SALAD) consistently improves recall, yielding 1-3% relative gains and up to 8% on challenging real flood imagery. By embedding urban perception principles into attention mechanisms, VPR-AttLLM bridges human-like spatial reasoning with modern VPR architectures. Its plug-and-play design and cross-source robustness offer a scalable solution for rapid geo-localization of crowdsourced crisis imagery, advancing cognitive urban resilience.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention
Xu, Fengyi
Ma, Jun
Qiu, Waishan
Guo, Cui
Cheng, Jack C. P.
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Crowdsourced social media imagery provides real-time visual evidence of urban flooding but often lacks reliable geographic metadata for emergency response. Existing Visual Place Recognition (VPR) models struggle to geo-localize these images due to cross-source domain shifts and visual distortions. We present VPR-AttLLM, a model-agnostic framework integrating the semantic reasoning and geospatial knowledge of Large Language Models (LLMs) into VPR pipelines via attention-guided descriptor enhancement. VPR-AttLLM uses LLMs to isolate location-informative regions and suppress transient noise, improving retrieval without model retraining or new data. We evaluate this framework across San Francisco and Hong Kong using established queries, synthetic flooding scenarios, and real social media flood images. Integrating VPR-AttLLM with state-of-the-art models (CosPlace, EigenPlaces, SALAD) consistently improves recall, yielding 1-3% relative gains and up to 8% on challenging real flood imagery. By embedding urban perception principles into attention mechanisms, VPR-AttLLM bridges human-like spatial reasoning with modern VPR architectures. Its plug-and-play design and cross-source robustness offer a scalable solution for rapid geo-localization of crowdsourced crisis imagery, advancing cognitive urban resilience.
title Enhancing Geo-localization for Crowdsourced Flood Imagery via LLM-Guided Attention
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2512.11811