Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908988002533376 |
|---|---|
| author | Yao, Siyuan Ghorbany, Siavash Ai, Kuangshi Cherukuthota, Arnav Forstchen, Meghan Korotasz, Alexis Sisk, Matthew Hu, Ming Wang, Chaoli |
| author_facet | Yao, Siyuan Ghorbany, Siavash Ai, Kuangshi Cherukuthota, Arnav Forstchen, Meghan Korotasz, Alexis Sisk, Matthew Hu, Ming Wang, Chaoli |
| contents | We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_21102 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery Yao, Siyuan Ghorbany, Siavash Ai, Kuangshi Cherukuthota, Arnav Forstchen, Meghan Korotasz, Alexis Sisk, Matthew Hu, Ming Wang, Chaoli Computer Vision and Pattern Recognition Artificial Intelligence We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort. |
| title | Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2604.21102 |