The Fantasia Bound on Constitutional Classifiers: Thermodynamic Limits of Jailbreak Defence

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteur principal: Eckert, Anthony
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866901971469860864
author Eckert, Anthony
author_facet Eckert, Anthony
contents Anthropic's constitutional classifiers (2026) withstood 3,000+ hours of red teaming with no universal jailbreak. We derive a thermodynamic upper bound on classifier effectiveness from the Fantasia Bound I(D;Y)+I(M;Y)≤H(Y). A classifier is a prohibition mechanism: its channel capacity sets the maximum Pe it can suppress. We prove that no single-channel classifier can defend against attacks above a critical Pe threshold Pe_c = exp(C/kT) where C is the classifier's information capacity. The prohibition-ritual pair architecture (two independent channels) raises the bound but does not eliminate it. Constitutional classifiers succeed because they approximate the two-channel architecture — the constitution provides the ritual (explicit reasoning about refusal), while the classifier provides the prohibition. Predictions: universal jailbreaks exist above Pe_c; defence requires increasing channel capacity, not classifier complexity.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19340888
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle The Fantasia Bound on Constitutional Classifiers: Thermodynamic Limits of Jailbreak Defence
Eckert, Anthony
constitutional classifiers
jailbreak defence
Fantasia Bound
conjugacy constraint
channel capacity
prohibition-ritual pair
Peclet number
thermodynamic bound
AI safety
red teaming
information theory
rate-distortion
Anthropic
classifier limits
Anthropic's constitutional classifiers (2026) withstood 3,000+ hours of red teaming with no universal jailbreak. We derive a thermodynamic upper bound on classifier effectiveness from the Fantasia Bound I(D;Y)+I(M;Y)≤H(Y). A classifier is a prohibition mechanism: its channel capacity sets the maximum Pe it can suppress. We prove that no single-channel classifier can defend against attacks above a critical Pe threshold Pe_c = exp(C/kT) where C is the classifier's information capacity. The prohibition-ritual pair architecture (two independent channels) raises the bound but does not eliminate it. Constitutional classifiers succeed because they approximate the two-channel architecture — the constitution provides the ritual (explicit reasoning about refusal), while the classifier provides the prohibition. Predictions: universal jailbreaks exist above Pe_c; defence requires increasing channel capacity, not classifier complexity.
title The Fantasia Bound on Constitutional Classifiers: Thermodynamic Limits of Jailbreak Defence
topic constitutional classifiers
jailbreak defence
Fantasia Bound
conjugacy constraint
channel capacity
prohibition-ritual pair
Peclet number
thermodynamic bound
AI safety
red teaming
information theory
rate-distortion
Anthropic
classifier limits
url https://doi.org/10.5281/zenodo.19340888