An Approach to Technical AGI Safety and Security

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Rohin, Irpan, Alex, Turner, Alexander Matt, Wang, Anna, Conmy, Arthur, Lindner, David, Brown-Cohen, Jonah, Ho, Lewis, Nanda, Neel, Popa, Raluca Ada, Jain, Rishub, Greig, Rory, Albanie, Samuel, Emmons, Scott, Farquhar, Sebastian, Krier, Sébastien, Rajamanoharan, Senthooran, Bridgers, Sophie, Ijitoye, Tobi, Everitt, Tom, Krakovna, Victoria, Varma, Vikrant, Mikulik, Vladimir, Kenton, Zachary, Orr, Dave, Legg, Shane, Goodman, Noah, Dafoe, Allan, Flynn, Four, Dragan, Anca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910901974597632
author Shah, Rohin
Irpan, Alex
Turner, Alexander Matt
Wang, Anna
Conmy, Arthur
Lindner, David
Brown-Cohen, Jonah
Ho, Lewis
Nanda, Neel
Popa, Raluca Ada
Jain, Rishub
Greig, Rory
Albanie, Samuel
Emmons, Scott
Farquhar, Sebastian
Krier, Sébastien
Rajamanoharan, Senthooran
Bridgers, Sophie
Ijitoye, Tobi
Everitt, Tom
Krakovna, Victoria
Varma, Vikrant
Mikulik, Vladimir
Kenton, Zachary
Orr, Dave
Legg, Shane
Goodman, Noah
Dafoe, Allan
Flynn, Four
Dragan, Anca
author_facet Shah, Rohin
Irpan, Alex
Turner, Alexander Matt
Wang, Anna
Conmy, Arthur
Lindner, David
Brown-Cohen, Jonah
Ho, Lewis
Nanda, Neel
Popa, Raluca Ada
Jain, Rishub
Greig, Rory
Albanie, Samuel
Emmons, Scott
Farquhar, Sebastian
Krier, Sébastien
Rajamanoharan, Senthooran
Bridgers, Sophie
Ijitoye, Tobi
Everitt, Tom
Krakovna, Victoria
Varma, Vikrant
Mikulik, Vladimir
Kenton, Zachary
Orr, Dave
Legg, Shane
Goodman, Noah
Dafoe, Allan
Flynn, Four
Dragan, Anca
contents Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough to significantly harm humanity. We identify four areas of risk: misuse, misalignment, mistakes, and structural risks. Of these, we focus on technical approaches to misuse and misalignment. For misuse, our strategy aims to prevent threat actors from accessing dangerous capabilities, by proactively identifying dangerous capabilities, and implementing robust security, access restrictions, monitoring, and model safety mitigations. To address misalignment, we outline two lines of defense. First, model-level mitigations such as amplified oversight and robust training can help to build an aligned model. Second, system-level security measures such as monitoring and access control can mitigate harm even if the model is misaligned. Techniques from interpretability, uncertainty estimation, and safer design patterns can enhance the effectiveness of these mitigations. Finally, we briefly outline how these ingredients could be combined to produce safety cases for AGI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_01849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Approach to Technical AGI Safety and Security
Shah, Rohin
Irpan, Alex
Turner, Alexander Matt
Wang, Anna
Conmy, Arthur
Lindner, David
Brown-Cohen, Jonah
Ho, Lewis
Nanda, Neel
Popa, Raluca Ada
Jain, Rishub
Greig, Rory
Albanie, Samuel
Emmons, Scott
Farquhar, Sebastian
Krier, Sébastien
Rajamanoharan, Senthooran
Bridgers, Sophie
Ijitoye, Tobi
Everitt, Tom
Krakovna, Victoria
Varma, Vikrant
Mikulik, Vladimir
Kenton, Zachary
Orr, Dave
Legg, Shane
Goodman, Noah
Dafoe, Allan
Flynn, Four
Dragan, Anca
Artificial Intelligence
Computers and Society
Machine Learning
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough to significantly harm humanity. We identify four areas of risk: misuse, misalignment, mistakes, and structural risks. Of these, we focus on technical approaches to misuse and misalignment. For misuse, our strategy aims to prevent threat actors from accessing dangerous capabilities, by proactively identifying dangerous capabilities, and implementing robust security, access restrictions, monitoring, and model safety mitigations. To address misalignment, we outline two lines of defense. First, model-level mitigations such as amplified oversight and robust training can help to build an aligned model. Second, system-level security measures such as monitoring and access control can mitigate harm even if the model is misaligned. Techniques from interpretability, uncertainty estimation, and safer design patterns can enhance the effectiveness of these mitigations. Finally, we briefly outline how these ingredients could be combined to produce safety cases for AGI systems.
title An Approach to Technical AGI Safety and Security
topic Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2504.01849