On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ball, Sarah, Gluch, Greg, Goldwasser, Shafi, Kreuter, Frauke, Reingold, Omer, Rothblum, Guy N.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911048504705024
author Ball, Sarah
Gluch, Greg
Goldwasser, Shafi
Kreuter, Frauke
Reingold, Omer
Rothblum, Guy N.
author_facet Ball, Sarah
Gluch, Greg
Goldwasser, Shafi
Kreuter, Frauke
Reingold, Omer
Rothblum, Guy N.
contents With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering of the input prompt before it reaches the model, and filtering the output after generation. Our main results demonstrate computational challenges in filtering both prompts and outputs. First, we show that there exist LLMs for which there are no efficient prompt filters: adversarial prompts that elicit harmful behavior can be easily constructed, which are computationally indistinguishable from benign prompts for any efficient filter. Our second main result identifies a natural setting in which output filtering is computationally intractable. All of our separation results are under cryptographic hardness assumptions. In addition to these core findings, we also formalize and study relaxed mitigation approaches, demonstrating further computational barriers. We conclude that safety cannot be achieved by designing filters external to the LLM internals (architecture and weights); in particular, black-box access to the LLM will not suffice. Based on our technical results, we argue that an aligned AI system's intelligence cannot be separated from its judgment.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07341
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
Ball, Sarah
Gluch, Greg
Goldwasser, Shafi
Kreuter, Frauke
Reingold, Omer
Rothblum, Guy N.
Artificial Intelligence
Cryptography and Security
With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering of the input prompt before it reaches the model, and filtering the output after generation. Our main results demonstrate computational challenges in filtering both prompts and outputs. First, we show that there exist LLMs for which there are no efficient prompt filters: adversarial prompts that elicit harmful behavior can be easily constructed, which are computationally indistinguishable from benign prompts for any efficient filter. Our second main result identifies a natural setting in which output filtering is computationally intractable. All of our separation results are under cryptographic hardness assumptions. In addition to these core findings, we also formalize and study relaxed mitigation approaches, demonstrating further computational barriers. We conclude that safety cannot be achieved by designing filters external to the LLM internals (architecture and weights); in particular, black-box access to the LLM will not suffice. Based on our technical results, we argue that an aligned AI system's intelligence cannot be separated from its judgment.
title On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2507.07341