Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lab, Shanghai AI, :, Chen, Xiaoyang, Chen, Yunhao, Chen, Zeren, Chen, Zhiyun, Cui, Hanyun, Duan, Yawen, Guo, Jiaxuan, Guo, Qi, Hu, Xuhao, Huang, Hong, Huang, Lige, Li, Chunxiao, Li, Juncheng, Lin, Qihao, Liu, Dongrui, Liu, Xinmin, Liu, Zicheng, Lu, Chaochao, Lu, Xiaoya, Qu, Jingjing, Ren, Qibing, Shao, Jing, Shi, Jingwei, Sun, Jingwei, Wang, Peng, Wang, Weibing, Xu, Jia, Yan, Lewen, Yu, Xiao, Yu, Yi, Zhang, Boxuan, Zhang, Jie, Zhang, Weichen, Zheng, Zhijie, Zhou, Tianyi, Zhou, Bowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916864610795520
author Lab, Shanghai AI
:
Chen, Xiaoyang
Chen, Yunhao
Chen, Zeren
Chen, Zhiyun
Cui, Hanyun
Duan, Yawen
Guo, Jiaxuan
Guo, Qi
Hu, Xuhao
Huang, Hong
Huang, Lige
Li, Chunxiao
Li, Juncheng
Lin, Qihao
Liu, Dongrui
Liu, Xinmin
Liu, Zicheng
Lu, Chaochao
Lu, Xiaoya
Qu, Jingjing
Ren, Qibing
Shao, Jing
Shi, Jingwei
Sun, Jingwei
Wang, Peng
Wang, Weibing
Xu, Jia
Yan, Lewen
Yu, Xiao
Yu, Yi
Zhang, Boxuan
Zhang, Jie
Zhang, Weichen
Zheng, Zhijie
Zhou, Tianyi
Zhou, Bowen
author_facet Lab, Shanghai AI
:
Chen, Xiaoyang
Chen, Yunhao
Chen, Zeren
Chen, Zhiyun
Cui, Hanyun
Duan, Yawen
Guo, Jiaxuan
Guo, Qi
Hu, Xuhao
Huang, Hong
Huang, Lige
Li, Chunxiao
Li, Juncheng
Lin, Qihao
Liu, Dongrui
Liu, Xinmin
Liu, Zicheng
Lu, Chaochao
Lu, Xiaoya
Qu, Jingjing
Ren, Qibing
Shao, Jing
Shi, Jingwei
Sun, Jingwei
Wang, Peng
Wang, Weibing
Xu, Jia
Yan, Lewen
Yu, Xiao
Yu, Yi
Zhang, Boxuan
Zhang, Jie
Zhang, Weichen
Zheng, Zhijie
Zhou, Tianyi
Zhou, Bowen
contents To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (deployment environment, threat source, enabling capability) from the Frontier AI Risk Management Framework (v1.0) (SafeWork-F1-Framework), we identify critical risks in seven areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled autonomous AI R\&D, strategic deception and scheming, self-replication, and collusion. Guided by the "AI-$45^\circ$ Law," we evaluate these risks using "red lines" (intolerable thresholds) and "yellow lines" (early warning indicators) to define risk zones: green (manageable risk for routine deployment and continuous monitoring), yellow (requiring strengthened mitigations and controlled deployment), and red (necessitating suspension of development and/or deployment). Experimental results show that all recent frontier AI models reside in green and yellow zones, without crossing red lines. Specifically, no evaluated models cross the yellow line for cyber offense or uncontrolled AI R\&D risks. For self-replication, and strategic deception and scheming, most models remain in the green zone, except for certain reasoning models in the yellow zone. In persuasion and manipulation, most models are in the yellow zone due to their effective influence on humans. For biological and chemical risks, we are unable to rule out the possibility of most models residing in the yellow zone, although detailed threat modeling and in-depth assessment are required to make further claims. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16534
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
Lab, Shanghai AI
:
Chen, Xiaoyang
Chen, Yunhao
Chen, Zeren
Chen, Zhiyun
Cui, Hanyun
Duan, Yawen
Guo, Jiaxuan
Guo, Qi
Hu, Xuhao
Huang, Hong
Huang, Lige
Li, Chunxiao
Li, Juncheng
Lin, Qihao
Liu, Dongrui
Liu, Xinmin
Liu, Zicheng
Lu, Chaochao
Lu, Xiaoya
Qu, Jingjing
Ren, Qibing
Shao, Jing
Shi, Jingwei
Sun, Jingwei
Wang, Peng
Wang, Weibing
Xu, Jia
Yan, Lewen
Yu, Xiao
Yu, Yi
Zhang, Boxuan
Zhang, Jie
Zhang, Weichen
Zheng, Zhijie
Zhou, Tianyi
Zhou, Bowen
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (deployment environment, threat source, enabling capability) from the Frontier AI Risk Management Framework (v1.0) (SafeWork-F1-Framework), we identify critical risks in seven areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled autonomous AI R\&D, strategic deception and scheming, self-replication, and collusion. Guided by the "AI-$45^\circ$ Law," we evaluate these risks using "red lines" (intolerable thresholds) and "yellow lines" (early warning indicators) to define risk zones: green (manageable risk for routine deployment and continuous monitoring), yellow (requiring strengthened mitigations and controlled deployment), and red (necessitating suspension of development and/or deployment). Experimental results show that all recent frontier AI models reside in green and yellow zones, without crossing red lines. Specifically, no evaluated models cross the yellow line for cyber offense or uncontrolled AI R\&D risks. For self-replication, and strategic deception and scheming, most models remain in the green zone, except for certain reasoning models in the yellow zone. In persuasion and manipulation, most models are in the yellow zone due to their effective influence on humans. For biological and chemical risks, we are unable to rule out the possibility of most models residing in the yellow zone, although detailed threat modeling and in-depth assessment are required to make further claims. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.
title Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.16534