STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917245097082880 |
|---|---|
| author | Barrett, Steve Bruvere, Anna Fillingham, Sean P. Rhodes, Catherine Vergani, Stefano |
| author_facet | Barrett, Steve Bruvere, Anna Fillingham, Sean P. Rhodes, Catherine Vergani, Stefano |
| contents | A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to future existential risks, and a range of loss of control pathways from rapid AI self-exfiltration scenarios to more gradual disempowerment scenarios. In this work we set out to firstly, provide a more structured framework for discussing and characterizing loss of control and secondly, to use this framework to assist those responsible for the safe operation of AI-containing socio-technical systems to identify causal factors leading to loss of control. We explore how these two needs can be better met by making use of a methodology developed within the safety-critical systems community known as STAMP and its associated hazard analysis technique of STPA. We select the STAMP methodology primarily because it is based around a world-view that socio-technical systems can be functionally modeled as control structures, and that safety issues arise when there is a loss of control in these structures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_17600 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems Barrett, Steve Bruvere, Anna Fillingham, Sean P. Rhodes, Catherine Vergani, Stefano Computers and Society A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to future existential risks, and a range of loss of control pathways from rapid AI self-exfiltration scenarios to more gradual disempowerment scenarios. In this work we set out to firstly, provide a more structured framework for discussing and characterizing loss of control and secondly, to use this framework to assist those responsible for the safe operation of AI-containing socio-technical systems to identify causal factors leading to loss of control. We explore how these two needs can be better met by making use of a methodology developed within the safety-critical systems community known as STAMP and its associated hazard analysis technique of STPA. We select the STAMP methodology primarily because it is based around a world-view that socio-technical systems can be functionally modeled as control structures, and that safety issues arise when there is a loss of control in these structures. |
| title | STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems |
| topic | Computers and Society |
| url | https://arxiv.org/abs/2512.17600 |