Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Donohue, Jared Joseph, Lamb, Kara D.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912634209566720
author Donohue, Jared Joseph
Lamb, Kara D.
author_facet Donohue, Jared Joseph
Lamb, Kara D.
contents Cloud seeding, a weather modification technique used to increase precipitation, has been practiced in the western United States since the 1940s. However, comprehensive datasets are not currently available to analyze these efforts. To address this gap, we present a structured dataset of reported cloud seeding activities in the U.S. from 2000-2025, including the project name, year, season, state, operator, seeding agent, apparatus used for deployment, stated purpose, target area, control area, start date, and end date. Combining our multi-stage PDF-to-text extraction pipeline with OpenAI's o3 large language model (LLM), we processed 832 historical reports from the National Oceanic and Atmospheric Administration (NOAA). The resulting dataset demonstrates 98.38% estimated accuracy, based on manual review of 200 randomly sampled records, and is publicly available on Zenodo. This dataset addresses the gap in cloud seeding data and demonstrates the potential for LLMs to extract structured information from historical environmental documents. More broadly, this work provides a scalable framework for unlocking historical data from scanned documents across scientific domains.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01555
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM
Donohue, Jared Joseph
Lamb, Kara D.
Atmospheric and Oceanic Physics
Cloud seeding, a weather modification technique used to increase precipitation, has been practiced in the western United States since the 1940s. However, comprehensive datasets are not currently available to analyze these efforts. To address this gap, we present a structured dataset of reported cloud seeding activities in the U.S. from 2000-2025, including the project name, year, season, state, operator, seeding agent, apparatus used for deployment, stated purpose, target area, control area, start date, and end date. Combining our multi-stage PDF-to-text extraction pipeline with OpenAI's o3 large language model (LLM), we processed 832 historical reports from the National Oceanic and Atmospheric Administration (NOAA). The resulting dataset demonstrates 98.38% estimated accuracy, based on manual review of 200 randomly sampled records, and is publicly available on Zenodo. This dataset addresses the gap in cloud seeding data and demonstrates the potential for LLMs to extract structured information from historical environmental documents. More broadly, this work provides a scalable framework for unlocking historical data from scanned documents across scientific domains.
title Structured dataset of reported cloud seeding activities in the United States (2000-2025) using an LLM
topic Atmospheric and Oceanic Physics
url https://arxiv.org/abs/2505.01555