Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Torneiro, André, Monteiro, Diogo, Novais, Paulo, Henriques, Pedro Rangel, Rodrigues, Nuno F.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914091492179968
author Torneiro, André
Monteiro, Diogo
Novais, Paulo
Henriques, Pedro Rangel
Rodrigues, Nuno F.
author_facet Torneiro, André
Monteiro, Diogo
Novais, Paulo
Henriques, Pedro Rangel
Rodrigues, Nuno F.
contents Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved. Current state-of-the-art approaches typically rely on a combination of IoT sensors and manual inspections, which are costly, difficult to scale, and often misaligned with citizens' perception formed through direct visual observation. This raises a critical question: Can machines now "see" like citizens and infer informed opinions about the condition of urban infrastructure? Vision-Language Models (VLMs), which integrate visual understanding with natural language reasoning, have recently demonstrated impressive capabilities in processing complex visual information, turning them into a promising technology to address this challenge. This systematic review investigates the role of VLMs in urban monitoring, with particular emphasis on zero-shot applications. Following the PRISMA methodology, we analyzed 32 peer-reviewed studies published between 2021 and 2025 to address four core research questions: (1) What urban monitoring tasks have been effectively addressed using VLMs? (2) Which VLM architectures and frameworks are most commonly used and demonstrate superior performance? (3) What datasets and resources support this emerging field? (4) How are VLM-based applications evaluated, and what performance levels have been reported?
format Preprint
id arxiv_https___arxiv_org_abs_2510_12400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
Torneiro, André
Monteiro, Diogo
Novais, Paulo
Henriques, Pedro Rangel
Rodrigues, Nuno F.
Computer Vision and Pattern Recognition
Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved. Current state-of-the-art approaches typically rely on a combination of IoT sensors and manual inspections, which are costly, difficult to scale, and often misaligned with citizens' perception formed through direct visual observation. This raises a critical question: Can machines now "see" like citizens and infer informed opinions about the condition of urban infrastructure? Vision-Language Models (VLMs), which integrate visual understanding with natural language reasoning, have recently demonstrated impressive capabilities in processing complex visual information, turning them into a promising technology to address this challenge. This systematic review investigates the role of VLMs in urban monitoring, with particular emphasis on zero-shot applications. Following the PRISMA methodology, we analyzed 32 peer-reviewed studies published between 2021 and 2025 to address four core research questions: (1) What urban monitoring tasks have been effectively addressed using VLMs? (2) Which VLM architectures and frameworks are most commonly used and demonstrate superior performance? (3) What datasets and resources support this emerging field? (4) How are VLM-based applications evaluated, and what performance levels have been reported?
title Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12400