Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pokharel, Gaurab, Farabi, Shafkat, Fowler, Patrick J., Das, Sanmay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914021464080384
author Pokharel, Gaurab
Farabi, Shafkat
Fowler, Patrick J.
Das, Sanmay
author_facet Pokharel, Gaurab
Farabi, Shafkat
Fowler, Patrick J.
Das, Sanmay
contents A surge of recent work explores the ethical and societal implications of large-scale AI models that make "moral" judgments. Much of this literature focuses either on alignment with human judgments through various thought experiments or on the group fairness implications of AI judgments. However, the most immediate and likely use of AI is to help or fully replace the so-called street-level bureaucrats, the individuals deciding to allocate scarce social resources or approve benefits. There is a rich history underlying how principles of local justice determine how society decides on prioritization mechanisms in such domains. In this paper, we examine how well LLM judgments align with human judgments, as well as with socially and politically determined vulnerability scoring systems currently used in the domain of homelessness resource allocation. Crucially, we use real data on those needing services (maintaining strict confidentiality by only using local large models) to perform our analyses. We find that LLM prioritizations are extremely inconsistent in several ways: internally on different runs, between different LLMs, and between LLMs and the vulnerability scoring systems. At the same time, LLMs demonstrate qualitative consistency with lay human judgments in pairwise testing. Findings call into question the readiness of current generation AI systems for naive integration in high-stakes societal decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
Pokharel, Gaurab
Farabi, Shafkat
Fowler, Patrick J.
Das, Sanmay
Computers and Society
Artificial Intelligence
A surge of recent work explores the ethical and societal implications of large-scale AI models that make "moral" judgments. Much of this literature focuses either on alignment with human judgments through various thought experiments or on the group fairness implications of AI judgments. However, the most immediate and likely use of AI is to help or fully replace the so-called street-level bureaucrats, the individuals deciding to allocate scarce social resources or approve benefits. There is a rich history underlying how principles of local justice determine how society decides on prioritization mechanisms in such domains. In this paper, we examine how well LLM judgments align with human judgments, as well as with socially and politically determined vulnerability scoring systems currently used in the domain of homelessness resource allocation. Crucially, we use real data on those needing services (maintaining strict confidentiality by only using local large models) to perform our analyses. We find that LLM prioritizations are extremely inconsistent in several ways: internally on different runs, between different LLMs, and between LLMs and the vulnerability scoring systems. At the same time, LLMs demonstrate qualitative consistency with lay human judgments in pairwise testing. Findings call into question the readiness of current generation AI systems for naive integration in high-stakes societal decision-making.
title Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
topic Computers and Society
Artificial Intelligence
url https://arxiv.org/abs/2508.08193