Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yiming, Gupta, Archit, Esparza, Miguel, Ho, Yu-Hsuan, Sebastian, Antonia, Weas, Hannah, Houck, Rose, Mostafavi, Ali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910010766786560
author Xiao, Yiming
Gupta, Archit
Esparza, Miguel
Ho, Yu-Hsuan
Sebastian, Antonia
Weas, Hannah
Houck, Rose
Mostafavi, Ali
author_facet Xiao, Yiming
Gupta, Archit
Esparza, Miguel
Ho, Yu-Hsuan
Sebastian, Antonia
Weas, Hannah
Houck, Rose
Mostafavi, Ali
contents Building-level occupancy after disasters is vital for triage, inspections, utility re-energization, and equitable resource allocation. Overhead imagery provides rapid coverage but often misses facade and access cues that determine habitability, while street-view imagery captures those details but is sparse and difficult to align with parcels. We present FacadeTrack, a street-level, language-guided framework that links panoramic video to parcels, rectifies views to facades, and elicits interpretable attributes (for example, entry blockage, temporary coverings, localized debris) that drive two decision strategies: a transparent one-stage rule and a two-stage design that separates perception from conservative reasoning. Evaluated across two post-Hurricane Helene surveys, the two-stage approach achieves a precision of 0.927, a recall of 0.781, and an F-1 score of 0.848, compared with the one-stage baseline at a precision of 0.943, a recall of 0.728, and an F-1 score of 0.822. Beyond accuracy, intermediate attributes and spatial diagnostics reveal where and why residual errors occur, enabling targeted quality control. The pipeline provides auditable, scalable occupancy assessments suitable for integration into geospatial and emergency-management workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20628
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
Xiao, Yiming
Gupta, Archit
Esparza, Miguel
Ho, Yu-Hsuan
Sebastian, Antonia
Weas, Hannah
Houck, Rose
Mostafavi, Ali
Computer Vision and Pattern Recognition
Building-level occupancy after disasters is vital for triage, inspections, utility re-energization, and equitable resource allocation. Overhead imagery provides rapid coverage but often misses facade and access cues that determine habitability, while street-view imagery captures those details but is sparse and difficult to align with parcels. We present FacadeTrack, a street-level, language-guided framework that links panoramic video to parcels, rectifies views to facades, and elicits interpretable attributes (for example, entry blockage, temporary coverings, localized debris) that drive two decision strategies: a transparent one-stage rule and a two-stage design that separates perception from conservative reasoning. Evaluated across two post-Hurricane Helene surveys, the two-stage approach achieves a precision of 0.927, a recall of 0.781, and an F-1 score of 0.848, compared with the one-stage baseline at a precision of 0.943, a recall of 0.728, and an F-1 score of 0.822. Beyond accuracy, intermediate attributes and spatial diagnostics reveal where and why residual errors occur, enabling targeted quality control. The pipeline provides auditable, scalable occupancy assessments suitable for integration into geospatial and emergency-management workflows.
title Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.20628