Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Siyuan, Ghorbany, Siavash, Ai, Kuangshi, Cherukuthota, Arnav, Forstchen, Meghan, Korotasz, Alexis, Sisk, Matthew, Hu, Ming, Wang, Chaoli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908988002533376
author Yao, Siyuan
Ghorbany, Siavash
Ai, Kuangshi
Cherukuthota, Arnav
Forstchen, Meghan
Korotasz, Alexis
Sisk, Matthew
Hu, Ming
Wang, Chaoli
author_facet Yao, Siyuan
Ghorbany, Siavash
Ai, Kuangshi
Cherukuthota, Arnav
Forstchen, Meghan
Korotasz, Alexis
Sisk, Matthew
Hu, Ming
Wang, Chaoli
contents We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
Yao, Siyuan
Ghorbany, Siavash
Ai, Kuangshi
Cherukuthota, Arnav
Forstchen, Meghan
Korotasz, Alexis
Sisk, Matthew
Hu, Ming
Wang, Chaoli
Computer Vision and Pattern Recognition
Artificial Intelligence
We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort.
title Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.21102