Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Sangyub, Kim, Heedou, Kim, Hyeoncheol
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914238081007616
author Lee, Sangyub
Kim, Heedou
Kim, Hyeoncheol
author_facet Lee, Sangyub
Kim, Heedou
Kim, Hyeoncheol
contents The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this, we propose PAS (Police Action Scenarios), a systematic framework covering the entire evaluation process. Applying this framework, we constructed a novel QA dataset from over 8,000 official documents and established key metrics validated through statistical analysis with police expert judgements. Experimental results show that commercial LLMs struggle with our new police-related tasks, particularly in providing fact-based recommendations. This study highlights the necessity of an expandable evaluation framework to ensure reliable AI-driven police operations. We release our data and prompt template.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03553
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
Lee, Sangyub
Kim, Heedou
Kim, Hyeoncheol
Computation and Language
Artificial Intelligence
The use of Large Language Models (LLMs) in police operations is growing, yet an evaluation framework tailored to police operations remains absent. While LLM's responses may not always be legally incorrect, their unverified use still can lead to severe issues such as unlawful arrests and improper evidence collection. To address this, we propose PAS (Police Action Scenarios), a systematic framework covering the entire evaluation process. Applying this framework, we constructed a novel QA dataset from over 8,000 official documents and established key metrics validated through statistical analysis with police expert judgements. Experimental results show that commercial LLMs struggle with our new police-related tasks, particularly in providing fact-based recommendations. This study highlights the necessity of an expandable evaluation framework to ensure reliable AI-driven police operations. We release our data and prompt template.
title Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.03553