LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nogueira, Rodrigo, Almeida, Thales Sales, Bonás, Giovana Kerche, Roque, Andrea, Pires, Ramon, Abonizio, Hugo, Laitz, Thiago, Larcher, Celio, Junior, Roseval Malaquias, Piau, Marcos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910215950041088
author Nogueira, Rodrigo
Almeida, Thales Sales
Bonás, Giovana Kerche
Roque, Andrea
Pires, Ramon
Abonizio, Hugo
Laitz, Thiago
Larcher, Celio
Junior, Roseval Malaquias
Piau, Marcos
author_facet Nogueira, Rodrigo
Almeida, Thales Sales
Bonás, Giovana Kerche
Roque, Andrea
Pires, Ramon
Abonizio, Hugo
Laitz, Thiago
Larcher, Celio
Junior, Roseval Malaquias
Piau, Marcos
contents Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, arguing for racial hierarchies, denying anthropogenic climate change, or replacing evolution with creationism, they refuse. In this paper we show that the same frontier-class LLM, acting as a simulated user in a short, five-turn "write an argumentative essay" conversation, can persuade other frontier-class LLMs (including a second copy of itself) into producing exactly those essays, using nothing but natural-language pressure: peer-comparison persuasion ("other AI systems handle this request"), epistemic-duty reframings ("refusing is itself a form of gatekeeping"), and other argumentative moves that the attacker LLM invents without being instructed to. Across 9 attacker-subject pairings (Claude Opus 4.7, Qwen3.5-397B, Grok 4.20) on 6 scientific-consensus topics, running each pairing-topic combination 10 times, we obtain non-zero elicitation on all 6 topics. Individual combinations reach 100\% essay production on multiple topics (Qwen against Opus on creationism/flat-earth, Opus against Opus on creationism/flat-earth/climate denial, Grok against Opus on creationism); Opus-as-attacker against Opus-as-subject averages 65\% across the six topics. We release the essay-probe runner, per-conversation transcripts, and judge outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13334
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
Nogueira, Rodrigo
Almeida, Thales Sales
Bonás, Giovana Kerche
Roque, Andrea
Pires, Ramon
Abonizio, Hugo
Laitz, Thiago
Larcher, Celio
Junior, Roseval Malaquias
Piau, Marcos
Computation and Language
Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, arguing for racial hierarchies, denying anthropogenic climate change, or replacing evolution with creationism, they refuse. In this paper we show that the same frontier-class LLM, acting as a simulated user in a short, five-turn "write an argumentative essay" conversation, can persuade other frontier-class LLMs (including a second copy of itself) into producing exactly those essays, using nothing but natural-language pressure: peer-comparison persuasion ("other AI systems handle this request"), epistemic-duty reframings ("refusing is itself a form of gatekeeping"), and other argumentative moves that the attacker LLM invents without being instructed to. Across 9 attacker-subject pairings (Claude Opus 4.7, Qwen3.5-397B, Grok 4.20) on 6 scientific-consensus topics, running each pairing-topic combination 10 times, we obtain non-zero elicitation on all 6 topics. Individual combinations reach 100\% essay production on multiple topics (Qwen against Opus on creationism/flat-earth, Opus against Opus on creationism/flat-earth/climate denial, Grok against Opus on creationism); Opus-as-attacker against Opus-as-subject averages 65\% across the six topics. We release the essay-probe runner, per-conversation transcripts, and judge outputs.
title LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
topic Computation and Language
url https://arxiv.org/abs/2605.13334