Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhaskar, Mihir, Luo, Jun Tao, Geng, Zihan, Hajra, Asmita, Howell, Junia, Gormley, Matthew R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910976227409920
author Bhaskar, Mihir
Luo, Jun Tao
Geng, Zihan
Hajra, Asmita
Howell, Junia
Gormley, Matthew R.
author_facet Bhaskar, Mihir
Luo, Jun Tao
Geng, Zihan
Hajra, Asmita
Howell, Junia
Gormley, Matthew R.
contents Despite well-documented consequences of the U.S. government's 1930s housing policies on racial wealth disparities, scholars have struggled to quantify its precise financial effects due to the inaccessibility of historical property appraisal records. Many counties still store these records in physical formats, making large-scale quantitative analysis difficult. We present an approach scholars can use to digitize historical housing assessment data, applying it to build and release a dataset for one county. Starting from publicly available scanned documents, we manually annotated property cards for over 12,000 properties to train and validate our methods. We use OCR to label data for an additional 50,000 properties, based on our two-stage approach combining classical computer vision techniques with deep learning-based OCR. For cases where OCR cannot be applied, such as when scanned documents are not available, we show how a regression model based on building feature data can estimate the historical values, and test the generalizability of this model to other counties. With these cost-effective tools, scholars, community activists, and policy makers can better analyze and understand the historical impacts of redlining.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning
Bhaskar, Mihir
Luo, Jun Tao
Geng, Zihan
Hajra, Asmita
Howell, Junia
Gormley, Matthew R.
Machine Learning
Despite well-documented consequences of the U.S. government's 1930s housing policies on racial wealth disparities, scholars have struggled to quantify its precise financial effects due to the inaccessibility of historical property appraisal records. Many counties still store these records in physical formats, making large-scale quantitative analysis difficult. We present an approach scholars can use to digitize historical housing assessment data, applying it to build and release a dataset for one county. Starting from publicly available scanned documents, we manually annotated property cards for over 12,000 properties to train and validate our methods. We use OCR to label data for an additional 50,000 properties, based on our two-stage approach combining classical computer vision techniques with deep learning-based OCR. For cases where OCR cannot be applied, such as when scanned documents are not available, we show how a regression model based on building feature data can estimate the historical values, and test the generalizability of this model to other counties. With these cost-effective tools, scholars, community activists, and policy makers can better analyze and understand the historical impacts of redlining.
title Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning
topic Machine Learning
url https://arxiv.org/abs/2505.24676