SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barbieri, Sidnei, de Meneses, Leonardo Vaz, Ferraz, Ágney Lopes Roth, Júnior, Lourenço Alves Pereira
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914536755298304
author Barbieri, Sidnei
de Meneses, Leonardo Vaz
Ferraz, Ágney Lopes Roth
Júnior, Lourenço Alves Pereira
author_facet Barbieri, Sidnei
de Meneses, Leonardo Vaz
Ferraz, Ágney Lopes Roth
Júnior, Lourenço Alves Pereira
contents Security operations centers (SOCs) are beginning to use large language models (LLMs) as copilots to draft incident-response plans. These plans may include actions that are valid per the catalog but still violate mandatory steps, required ordering, or approval gates before analyst review. SOCpilot makes this compliance question measurable at the plan boundary. It fixes the incident package, action catalog, policy rules, verifier, and public evidence surface. Next, it verifies the copilot's proposed action trace. We evaluate two LLM providers on 200 real incidents from an anonymized production SOC in a financial-sector case study. We compare their plans to paired analyst-authored references from the same security orchestration, automation, and response (SOAR) cases. An identical inline policy text moves the two providers in opposite directions. A deterministic verifier removes 466 non-compliant, approval-gated actions, without reducing baseline-task recall. Aggregate rates remain stable across 3 reruns of the fixed corpus. The official evidence focuses on approval-gated decisions regarding recovery and containment. Separately, the artifact exposes zero-cost readiness checks for mandatory and ordering repairs. We release the runnable artifact so independent reviewers can rederive the public results without access to private incident data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05501
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response
Barbieri, Sidnei
de Meneses, Leonardo Vaz
Ferraz, Ágney Lopes Roth
Júnior, Lourenço Alves Pereira
Cryptography and Security
Security operations centers (SOCs) are beginning to use large language models (LLMs) as copilots to draft incident-response plans. These plans may include actions that are valid per the catalog but still violate mandatory steps, required ordering, or approval gates before analyst review. SOCpilot makes this compliance question measurable at the plan boundary. It fixes the incident package, action catalog, policy rules, verifier, and public evidence surface. Next, it verifies the copilot's proposed action trace. We evaluate two LLM providers on 200 real incidents from an anonymized production SOC in a financial-sector case study. We compare their plans to paired analyst-authored references from the same security orchestration, automation, and response (SOAR) cases. An identical inline policy text moves the two providers in opposite directions. A deterministic verifier removes 466 non-compliant, approval-gated actions, without reducing baseline-task recall. Aggregate rates remain stable across 3 reruns of the fixed corpus. The official evidence focuses on approval-gated decisions regarding recovery and containment. Separately, the artifact exposes zero-cost readiness checks for mandatory and ordering repairs. We release the runnable artifact so independent reviewers can rederive the public results without access to private incident data.
title SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response
topic Cryptography and Security
url https://arxiv.org/abs/2605.05501