Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dassanayake, Rishane, Demetroudi, Mario, Walpole, James, Lentati, Lindley, Brown, Jason R., Young, Edward James
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911061271117824
author Dassanayake, Rishane
Demetroudi, Mario
Walpole, James
Lentati, Lindley
Brown, Jason R.
Young, Edward James
author_facet Dassanayake, Rishane
Demetroudi, Mario
Walpole, James
Lentati, Lindley
Brown, Jason R.
Young, Edward James
contents Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are often the weakest link in cybersecurity systems, and a misaligned AI system deployed internally within a frontier company may seek to undermine human oversight by manipulating employees. Despite this growing threat, manipulation attacks have received little attention, and no systematic framework exists for assessing and mitigating these risks. To address this, we provide a detailed explanation of why manipulation attacks are a significant threat and could lead to catastrophic outcomes. Additionally, we present a safety case framework for manipulation risk, structured around three core lines of argument: inability, control, and trustworthiness. For each argument, we specify evidence requirements, evaluation methodologies, and implementation considerations for direct application by AI companies. This paper provides the first systematic methodology for integrating manipulation risk into AI safety governance, offering AI companies a concrete foundation to assess and mitigate these threats before deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12872
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
Dassanayake, Rishane
Demetroudi, Mario
Walpole, James
Lentati, Lindley
Brown, Jason R.
Young, Edward James
Artificial Intelligence
Cryptography and Security
Human-Computer Interaction
Frontier AI systems are rapidly advancing in their capabilities to persuade, deceive, and influence human behaviour, with current models already demonstrating human-level persuasion and strategic deception in specific contexts. Humans are often the weakest link in cybersecurity systems, and a misaligned AI system deployed internally within a frontier company may seek to undermine human oversight by manipulating employees. Despite this growing threat, manipulation attacks have received little attention, and no systematic framework exists for assessing and mitigating these risks. To address this, we provide a detailed explanation of why manipulation attacks are a significant threat and could lead to catastrophic outcomes. Additionally, we present a safety case framework for manipulation risk, structured around three core lines of argument: inability, control, and trustworthiness. For each argument, we specify evidence requirements, evaluation methodologies, and implementation considerations for direct application by AI companies. This paper provides the first systematic methodology for integrating manipulation risk into AI safety governance, offering AI companies a concrete foundation to assess and mitigate these threats before deployment.
title Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
topic Artificial Intelligence
Cryptography and Security
Human-Computer Interaction
url https://arxiv.org/abs/2507.12872