Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917536096845824 |
|---|---|
| author | Yu, Zhe Xing, Wenpeng Li, Gaolei Xiong, Shuguang Wang, Hongzhi Teng, Xuyang Han, Meng |
| author_facet | Yu, Zhe Xing, Wenpeng Li, Gaolei Xiong, Shuguang Wang, Hongzhi Teng, Xuyang Han, Meng |
| contents | Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control gap -- they can detect contradictions in retrieved evidence yet still act on poisoned claims. We introduce the Cordon Principle -- no agent capable of final synthesis may access untrusted natural-language evidence -- and realize it through CORDON-MAS, a compartmentalized framework that enforces this principle architecturally by separating evidence extraction, cross-source audit, and answer synthesis into agents with asymmetric memory privileges. Across five BEIR datasets, CORDON-MAS reduces attack success rate by 92.4\% relative to undefended RAG. This reframes RAG poisoning from a detection problem to an information-flow control problem. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26754 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control Yu, Zhe Xing, Wenpeng Li, Gaolei Xiong, Shuguang Wang, Hongzhi Teng, Xuyang Han, Meng Cryptography and Security Artificial Intelligence Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control gap -- they can detect contradictions in retrieved evidence yet still act on poisoned claims. We introduce the Cordon Principle -- no agent capable of final synthesis may access untrusted natural-language evidence -- and realize it through CORDON-MAS, a compartmentalized framework that enforces this principle architecturally by separating evidence extraction, cross-source audit, and answer synthesis into agents with asymmetric memory privileges. Across five BEIR datasets, CORDON-MAS reduces attack success rate by 92.4\% relative to undefended RAG. This reframes RAG poisoning from a detection problem to an information-flow control problem. |
| title | Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2605.26754 |