"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karr Jr., Jonathan A., Herbst, Benjamin F., Sisk, Matthew L., Li, Xueyun, Hua, Ting, Hauenstein, Matthew, Curto, Georgina, Chawla, Nitesh V.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917231948988416
author Karr Jr., Jonathan A.
Herbst, Benjamin F.
Sisk, Matthew L.
Li, Xueyun
Hua, Ting
Hauenstein, Matthew
Curto, Georgina
Chawla, Nitesh V.
author_facet Karr Jr., Jonathan A.
Herbst, Benjamin F.
Sisk, Matthew L.
Li, Xueyun
Hua, Ting
Hauenstein, Matthew
Curto, Georgina
Chawla, Nitesh V.
contents Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experienced homelessness (PEH) in the U.S. in 2025. Social bias is a significant barrier to alleviation, shaping public perception and influencing policymaking. Given that online textual media and offline city council discourse reflect and influence part of public opinion, it provides valuable insights to identify and track social biases against PEH. We present a new, manually-annotated multi-domain dataset compiled from Reddit, X (formerly Twitter), news articles, and city council meeting minutes across ten U.S. cities. Our 16-category multi-label taxonomy creates a challenging long-tail classification problem: some categories appear in less than 1% of samples, while others exceed 70%. We find that small human-annotated datasets (1,702 samples) are insufficient for training effective classifiers, whether used to fine-tune encoder models or as few-shot examples for LLMs. To address this, we use GPT-4.1 to generate pseudo-labels on a larger unlabeled corpus. Training on this expanded dataset enables even small encoder models (ModernBERT, 150M parameters) to achieve 35.23 macro-F1, approaching GPT-4.1's 41.57. This demonstrates that \textbf{data quantity matters more than model size}, enabling low-cost, privacy-preserving deployment without relying on commercial APIs. Our results reveal that negative bias against PEH is prevalent both offline and online (especially on Reddit), with "not in my backyard" narratives showing the highest engagement. These findings uncover a type of ostracism that directly impacts poverty-reduction policymaking and provide actionable insights for practitioners addressing homelessness.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13187
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle "Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
Karr Jr., Jonathan A.
Herbst, Benjamin F.
Sisk, Matthew L.
Li, Xueyun
Hua, Ting
Hauenstein, Matthew
Curto, Georgina
Chawla, Nitesh V.
Computers and Society
Artificial Intelligence
Computation and Language
Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experienced homelessness (PEH) in the U.S. in 2025. Social bias is a significant barrier to alleviation, shaping public perception and influencing policymaking. Given that online textual media and offline city council discourse reflect and influence part of public opinion, it provides valuable insights to identify and track social biases against PEH. We present a new, manually-annotated multi-domain dataset compiled from Reddit, X (formerly Twitter), news articles, and city council meeting minutes across ten U.S. cities. Our 16-category multi-label taxonomy creates a challenging long-tail classification problem: some categories appear in less than 1% of samples, while others exceed 70%. We find that small human-annotated datasets (1,702 samples) are insufficient for training effective classifiers, whether used to fine-tune encoder models or as few-shot examples for LLMs. To address this, we use GPT-4.1 to generate pseudo-labels on a larger unlabeled corpus. Training on this expanded dataset enables even small encoder models (ModernBERT, 150M parameters) to achieve 35.23 macro-F1, approaching GPT-4.1's 41.57. This demonstrates that \textbf{data quantity matters more than model size}, enabling low-cost, privacy-preserving deployment without relying on commercial APIs. Our results reveal that negative bias against PEH is prevalent both offline and online (especially on Reddit), with "not in my backyard" narratives showing the highest engagement. These findings uncover a type of ostracism that directly impacts poverty-reduction policymaking and provide actionable insights for practitioners addressing homelessness.
title "Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.13187