Efficiency at Risk: Entropy-Induced Failure in Learned Decision Policies under Partial Observability

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Pérez Contreras, Benjamín Felipe
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901868632866816
author Pérez Contreras, Benjamín Felipe
author_facet Pérez Contreras, Benjamín Felipe
contents <p>This work presents an empirical evaluation of automated security probing policies operating under partial observability. We quantify efficiency gains achieved by learned decision models and identify entropy-induced failure regimes that emerge under increasing observational noise. By modeling the interaction as a Partially Observable Markov Decision Process (POMDP), we demonstrate that while learned policies significantly reduce average Time-to-Compromise, they exhibit heavy-tailed operational risk characterized by instability, action thrashing, and elevated detection likelihood. We further propose a hybrid meta-control architecture that bounds worst-case behavior by supervising learned policies with entropy-aware fallback mechanisms. Our results highlight a fundamental efficiency–stability trade-off and provide actionable guidance for the safe deployment of autonomous security decision systems in stochastic environments. -  </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19058748
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Efficiency at Risk: Entropy-Induced Failure in Learned Decision Policies under Partial Observability
Pérez Contreras, Benjamín Felipe
Security, Offensive Security, POMDP, Machine Learning, Penetration Testing, Decision Policies
<p>This work presents an empirical evaluation of automated security probing policies operating under partial observability. We quantify efficiency gains achieved by learned decision models and identify entropy-induced failure regimes that emerge under increasing observational noise. By modeling the interaction as a Partially Observable Markov Decision Process (POMDP), we demonstrate that while learned policies significantly reduce average Time-to-Compromise, they exhibit heavy-tailed operational risk characterized by instability, action thrashing, and elevated detection likelihood. We further propose a hybrid meta-control architecture that bounds worst-case behavior by supervising learned policies with entropy-aware fallback mechanisms. Our results highlight a fundamental efficiency–stability trade-off and provide actionable guidance for the safe deployment of autonomous security decision systems in stochastic environments. -  </p>
title Efficiency at Risk: Entropy-Induced Failure in Learned Decision Policies under Partial Observability
topic Security, Offensive Security, POMDP, Machine Learning, Penetration Testing, Decision Policies
url https://doi.org/10.5281/zenodo.19058748