Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mathew, Yohan, Matthews, Ollie, McCarthy, Robert, Velja, Joan, de Witt, Christian Schroeder, Cope, Dylan, Schoots, Nandi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914176445710336
author Mathew, Yohan
Matthews, Ollie
McCarthy, Robert
Velja, Joan
de Witt, Christian Schroeder
Cope, Dylan
Schoots, Nandi
author_facet Mathew, Yohan
Matthews, Ollie
McCarthy, Robert
Velja, Joan
de Witt, Christian Schroeder
Cope, Dylan
Schoots, Nandi
contents The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to the disadvantage of others has been identified as a central form of undesirable agent cooperation. The use of information hiding (steganography) in agent communications could render such collusion practically undetectable. This underscores the need for investigations into the possibility of such behaviours emerging and the robustness corresponding countermeasures. To investigate this problem we design two approaches -- a gradient-based reinforcement learning (GBRL) method and an in-context reinforcement learning (ICRL) method -- for reliably eliciting sophisticated LLM-generated linguistic text steganography. We demonstrate, for the first time, that unintended steganographic collusion in LLMs can arise due to mispecified reward incentives during training. Additionally, we find that standard mitigations -- both passive oversight of model outputs and active mitigation through communication paraphrasing -- are not fully effective at preventing this steganographic communication. Our findings imply that (i) emergence of steganographic collusion is a plausible concern that should be monitored and researched, and (ii) preventing emergence may require innovation in mitigation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03768
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
Mathew, Yohan
Matthews, Ollie
McCarthy, Robert
Velja, Joan
de Witt, Christian Schroeder
Cope, Dylan
Schoots, Nandi
Computation and Language
Cryptography and Security
Machine Learning
The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to the disadvantage of others has been identified as a central form of undesirable agent cooperation. The use of information hiding (steganography) in agent communications could render such collusion practically undetectable. This underscores the need for investigations into the possibility of such behaviours emerging and the robustness corresponding countermeasures. To investigate this problem we design two approaches -- a gradient-based reinforcement learning (GBRL) method and an in-context reinforcement learning (ICRL) method -- for reliably eliciting sophisticated LLM-generated linguistic text steganography. We demonstrate, for the first time, that unintended steganographic collusion in LLMs can arise due to mispecified reward incentives during training. Additionally, we find that standard mitigations -- both passive oversight of model outputs and active mitigation through communication paraphrasing -- are not fully effective at preventing this steganographic communication. Our findings imply that (i) emergence of steganographic collusion is a plausible concern that should be monitored and researched, and (ii) preventing emergence may require innovation in mitigation techniques.
title Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2410.03768