Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Zhe, Xing, Wenpeng, Li, Gaolei, Xiong, Shuguang, Wang, Hongzhi, Teng, Xuyang, Han, Meng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917536096845824
author Yu, Zhe
Xing, Wenpeng
Li, Gaolei
Xiong, Shuguang
Wang, Hongzhi
Teng, Xuyang
Han, Meng
author_facet Yu, Zhe
Xing, Wenpeng
Li, Gaolei
Xiong, Shuguang
Wang, Hongzhi
Teng, Xuyang
Han, Meng
contents Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control gap -- they can detect contradictions in retrieved evidence yet still act on poisoned claims. We introduce the Cordon Principle -- no agent capable of final synthesis may access untrusted natural-language evidence -- and realize it through CORDON-MAS, a compartmentalized framework that enforces this principle architecturally by separating evidence extraction, cross-source audit, and answer synthesis into agents with asymmetric memory privileges. Across five BEIR datasets, CORDON-MAS reduces attack success rate by 92.4\% relative to undefended RAG. This reframes RAG poisoning from a detection problem to an information-flow control problem.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26754
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
Yu, Zhe
Xing, Wenpeng
Li, Gaolei
Xiong, Shuguang
Wang, Hongzhi
Teng, Xuyang
Han, Meng
Cryptography and Security
Artificial Intelligence
Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control gap -- they can detect contradictions in retrieved evidence yet still act on poisoned claims. We introduce the Cordon Principle -- no agent capable of final synthesis may access untrusted natural-language evidence -- and realize it through CORDON-MAS, a compartmentalized framework that enforces this principle architecturally by separating evidence extraction, cross-source audit, and answer synthesis into agents with asymmetric memory privileges. Across five BEIR datasets, CORDON-MAS reduces attack success rate by 92.4\% relative to undefended RAG. This reframes RAG poisoning from a detection problem to an information-flow control problem.
title Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.26754