STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Barrett, Steve, Bruvere, Anna, Fillingham, Sean P., Rhodes, Catherine, Vergani, Stefano
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917245097082880
author Barrett, Steve
Bruvere, Anna
Fillingham, Sean P.
Rhodes, Catherine
Vergani, Stefano
author_facet Barrett, Steve
Bruvere, Anna
Fillingham, Sean P.
Rhodes, Catherine
Vergani, Stefano
contents A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to future existential risks, and a range of loss of control pathways from rapid AI self-exfiltration scenarios to more gradual disempowerment scenarios. In this work we set out to firstly, provide a more structured framework for discussing and characterizing loss of control and secondly, to use this framework to assist those responsible for the safe operation of AI-containing socio-technical systems to identify causal factors leading to loss of control. We explore how these two needs can be better met by making use of a methodology developed within the safety-critical systems community known as STAMP and its associated hazard analysis technique of STPA. We select the STAMP methodology primarily because it is based around a world-view that socio-technical systems can be functionally modeled as control structures, and that safety issues arise when there is a loss of control in these structures.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems
Barrett, Steve
Bruvere, Anna
Fillingham, Sean P.
Rhodes, Catherine
Vergani, Stefano
Computers and Society
A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to future existential risks, and a range of loss of control pathways from rapid AI self-exfiltration scenarios to more gradual disempowerment scenarios. In this work we set out to firstly, provide a more structured framework for discussing and characterizing loss of control and secondly, to use this framework to assist those responsible for the safe operation of AI-containing socio-technical systems to identify causal factors leading to loss of control. We explore how these two needs can be better met by making use of a methodology developed within the safety-critical systems community known as STAMP and its associated hazard analysis technique of STPA. We select the STAMP methodology primarily because it is based around a world-view that socio-technical systems can be functionally modeled as control structures, and that safety issues arise when there is a loss of control in these structures.
title STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems
topic Computers and Society
url https://arxiv.org/abs/2512.17600