Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Sangyub, Kim, Heedou, Kim, Hyeoncheol |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAPIS: Language Model-Augmented Police Investigation System
by: Kim, Heedou, et al.
Published: (2024)
by: Kim, Heedou, et al.
Published: (2024)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
by: Hwang, Yeonjun, et al.
Published: (2026)
by: Hwang, Yeonjun, et al.
Published: (2026)
Gender Bias in LLM-generated Interview Responses
by: Kong, Haein, et al.
Published: (2024)
by: Kong, Haein, et al.
Published: (2024)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Cognitive Bias in Decision-Making with LLMs
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
Ontology for Policing: Conceptual Knowledge Learning for Semantic Understanding and Reasoning in Law Enforcement Reports
by: Srbinovska, Anita, et al.
Published: (2026)
by: Srbinovska, Anita, et al.
Published: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Pedagogy-R1: Pedagogically-Aligned Reasoning Model with Balanced Educational Benchmark
by: Lee, Unggi, et al.
Published: (2025)
by: Lee, Unggi, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
by: Deng, Jiaqi, et al.
Published: (2025)
by: Deng, Jiaqi, et al.
Published: (2025)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
by: Kim, Byungjun, et al.
Published: (2024)
by: Kim, Byungjun, et al.
Published: (2024)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
KTCF: Actionable Recourse in Knowledge Tracing via Counterfactual Explanations for Education
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
LLMs for Explainable Business Decision-Making: A Reinforcement Learning Fine-Tuning Approach
by: Cheng, Xiang, et al.
Published: (2025)
by: Cheng, Xiang, et al.
Published: (2025)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
by: Lee, Hojae, et al.
Published: (2024)
by: Lee, Hojae, et al.
Published: (2024)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
by: Seo, Jean, et al.
Published: (2026)
by: Seo, Jean, et al.
Published: (2026)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
by: Lee, Kyeongryul, et al.
Published: (2025)
by: Lee, Kyeongryul, et al.
Published: (2025)
DEBATE: Devil's Advocate-Based Assessment and Text Evaluation
by: Kim, Alex, et al.
Published: (2024)
by: Kim, Alex, et al.
Published: (2024)
Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios
by: Xu, Shaochen, et al.
Published: (2024)
by: Xu, Shaochen, et al.
Published: (2024)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
by: Xiao, Yunpeng, et al.
Published: (2025)
by: Xiao, Yunpeng, et al.
Published: (2025)
LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education
by: Lee, Unggi, et al.
Published: (2024)
by: Lee, Unggi, et al.
Published: (2024)
GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
by: Kim, Kun-Woo, et al.
Published: (2025)
by: Kim, Kun-Woo, et al.
Published: (2025)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
by: Jain, Ojas, et al.
Published: (2026)
by: Jain, Ojas, et al.
Published: (2026)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
by: Kim, SungHo, et al.
Published: (2025)
by: Kim, SungHo, et al.
Published: (2025)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
SCRIPTMIND: Crime Script Inference and Cognitive Evaluation for LLM-based Social Engineering Scam Detection System
by: Kim, Heedou, et al.
Published: (2026)
by: Kim, Heedou, et al.
Published: (2026)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage
by: Srbinovska, Anita, et al.
Published: (2025)
by: Srbinovska, Anita, et al.
Published: (2025)
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
by: Sakabe, Ritsu, et al.
Published: (2025)
by: Sakabe, Ritsu, et al.
Published: (2025)
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
by: Lee, Yanggyu, et al.
Published: (2024)
by: Lee, Yanggyu, et al.
Published: (2024)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios
by: Yang, Tingyue, et al.
Published: (2025)
by: Yang, Tingyue, et al.
Published: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Similar Items
-
LAPIS: Language Model-Augmented Police Investigation System
by: Kim, Heedou, et al.
Published: (2024) -
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
by: Hwang, Yeonjun, et al.
Published: (2026) -
Gender Bias in LLM-generated Interview Responses
by: Kong, Haein, et al.
Published: (2024) -
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025) -
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)