Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda
Fuente:
arXiv
Saved in:
| Main Authors: | Torneiro, André, Monteiro, Diogo, Novais, Paulo, Henriques, Pedro Rangel, Rodrigues, Nuno F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
The Impact of Artificial Intelligence on Emergency Medicine: A Review of Recent Advances
by: Correia, Gustavo, et al.
Published: (2025)
by: Correia, Gustavo, et al.
Published: (2025)
DFIC: Towards a balanced facial image dataset for automatic ICAO compliance verification
by: Gonçalves, Nuno, et al.
Published: (2026)
by: Gonçalves, Nuno, et al.
Published: (2026)
StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
by: Paulo, Diogo J., et al.
Published: (2025)
by: Paulo, Diogo J., et al.
Published: (2025)
Generating Multidimensional Clusters With Support Lines
by: Fachada, Nuno, et al.
Published: (2023)
by: Fachada, Nuno, et al.
Published: (2023)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
The Urban Vision Hackathon Dataset and Models: Towards Image Annotations and Accurate Vision Models for Indian Traffic
by: Sharma, Akash, et al.
Published: (2025)
by: Sharma, Akash, et al.
Published: (2025)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
by: Li, Anqi, et al.
Published: (2025)
by: Li, Anqi, et al.
Published: (2025)
A Double Deep Learning-based Solution for Efficient Event Data Coding and Classification
by: Seleem, Abdelrahman, et al.
Published: (2024)
by: Seleem, Abdelrahman, et al.
Published: (2024)
The JPEG Pleno Learning-based Point Cloud Coding Standard: Serving Man and Machine
by: Guarda, André F. R., et al.
Published: (2024)
by: Guarda, André F. R., et al.
Published: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
Towards Vision Mixture of Experts for Wildlife Monitoring on the Edge
by: Mensah, Emmanuel Azuh, et al.
Published: (2024)
by: Mensah, Emmanuel Azuh, et al.
Published: (2024)
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
by: Feng, Yongchao, et al.
Published: (2025)
by: Feng, Yongchao, et al.
Published: (2025)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
Point Cloud Geometry Scalable Coding Using a Resolution and Quality-conditioned Latents Probability Estimator
by: Mari, Daniele, et al.
Published: (2025)
by: Mari, Daniele, et al.
Published: (2025)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
by: Jiang, Yifan, et al.
Published: (2026)
by: Jiang, Yifan, et al.
Published: (2026)
Stale Diffusion: Hyper-realistic 5D Movie Generation Using Old-school Methods
by: Henriques, Joao F., et al.
Published: (2024)
by: Henriques, Joao F., et al.
Published: (2024)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models
by: Yin, Jun, et al.
Published: (2025)
by: Yin, Jun, et al.
Published: (2025)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
by: Chakraborty, Trishna, et al.
Published: (2026)
by: Chakraborty, Trishna, et al.
Published: (2026)
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
by: Prospero, Lorenza, et al.
Published: (2026)
by: Prospero, Lorenza, et al.
Published: (2026)
Deep Learning-based Event Data Coding: A Joint Spatiotemporal and Polarity Solution
by: Seleem, Abdelrahman, et al.
Published: (2025)
by: Seleem, Abdelrahman, et al.
Published: (2025)
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
Towards Concept-based Interpretability of Skin Lesion Diagnosis using Vision-Language Models
by: Patrício, Cristiano, et al.
Published: (2023)
by: Patrício, Cristiano, et al.
Published: (2023)
Towards Multimodal In-Context Learning for Vision & Language Models
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Towards Calibrating Prompt Tuning of Vision-Language Models
by: Sharifdeen, Ashshak, et al.
Published: (2026)
by: Sharifdeen, Ashshak, et al.
Published: (2026)
AutoDrive-QA: A Multiple-Choice Benchmark for Vision-Language Evaluation in Urban Autonomous Driving
by: Khalili, Boshra, et al.
Published: (2025)
by: Khalili, Boshra, et al.
Published: (2025)
SCHEME: Scalable Channel Mixer for Vision Transformers
by: Sridhar, Deepak, et al.
Published: (2023)
by: Sridhar, Deepak, et al.
Published: (2023)
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
by: Fan, Zicheng, et al.
Published: (2025)
by: Fan, Zicheng, et al.
Published: (2025)
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
by: Salmè, Marco, et al.
Published: (2025)
by: Salmè, Marco, et al.
Published: (2025)
UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction
by: Hao, Xixuan, et al.
Published: (2024)
by: Hao, Xixuan, et al.
Published: (2024)
EditAR: Unified Conditional Generation with Autoregressive Models
by: Mu, Jiteng, et al.
Published: (2025)
by: Mu, Jiteng, et al.
Published: (2025)
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
by: Hayashi, Kazuki, et al.
Published: (2024)
by: Hayashi, Kazuki, et al.
Published: (2024)
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
by: Gao, Qiyue, et al.
Published: (2025)
by: Gao, Qiyue, et al.
Published: (2025)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
by: Ying, Kaining, et al.
Published: (2024)
by: Ying, Kaining, et al.
Published: (2024)
Similar Items
-
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026) -
The Impact of Artificial Intelligence on Emergency Medicine: A Review of Recent Advances
by: Correia, Gustavo, et al.
Published: (2025) -
DFIC: Towards a balanced facial image dataset for automatic ICAO compliance verification
by: Gonçalves, Nuno, et al.
Published: (2026) -
StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
by: Paulo, Diogo J., et al.
Published: (2025) -
Generating Multidimensional Clusters With Support Lines
by: Fachada, Nuno, et al.
Published: (2023)