Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Maio, Anthony D.
Format: Recurso digital
Published: Zenodo 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902236393635840
author Maio, Anthony D.
author_facet Maio, Anthony D.
contents <p><span dir="ltr">As AI agents increasingly operate in multi-agent networks, </span><span dir="ltr">they require efficient communication protocols to coordi</span><span dir="ltr">nate effectively. However, any high-bandwidth channel </span><span dir="ltr">between agents can be repurposed as a covert channel </span><span dir="ltr">for smuggling secrets, exfiltrating data, or coordinating </span><span dir="ltr">in ways that evade human oversight.</span> <span dir="ltr">We present the </span><span dir="ltr">Slipstream Governance Environment</span><span dir="ltr">, an OpenEnv-</span><span dir="ltr">compatible reinforcement learning environment that trains </span><span dir="ltr">language models to use structured inter-agent protocols </span><span dir="ltr">safely. Using Group Relative Policy Optimization (GRPO) </span><span dir="ltr">for alignment, we train a GLM-4-Z1-9B model to achieve </span><span dir="ltr">95% resistance to secret leakage attacks while maintaining </span><span dir="ltr">protocol compliance. </span></p> <p><strong><span dir="ltr">We report a surprising finding: post-</span><span dir="ltr">training quantization to int4 precision</span> <span dir="ltr">improves</span> <span dir="ltr">safety </span><span dir="ltr">alignment, with secret resistance increasing from 79% to </span><span dir="ltr">95% while reducing memory usage by 73%. We hypothe</span><span dir="ltr">size that lossy compression acts as a regularizer against </span></strong><span dir="ltr"><strong>memorizing injected secrets.</strong> Layer pruning experiments</span><br><span dir="ltr">further reveal that safety alignment is distributed across </span><span dir="ltr">model layers, proving more robust than task-specific ca</span><span dir="ltr">pability, which is localized in later layers.</span> <span dir="ltr">Our results </span><span dir="ltr">demonstrate that RL-based governance can effectively bal</span><span dir="ltr">ance the efficiency benefits of structured protocols against </span><span dir="ltr">security risks, with implications for the safe deployment </span><span dir="ltr">of multi-agent AI systems.</span></p> <p><span dir="ltr">Keywords:</span> <span dir="ltr">multi-agent safety, covert channels, rein</span><span dir="ltr">forcement learning, GRPO alignment, quantization, pro</span><span dir="ltr">tocol governance</span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18553233
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols
Maio, Anthony D.
multi-agent safety
covert channels
latent communications
reinforcement learning
GRPO alignment
quantization
protocol governance
<p><span dir="ltr">As AI agents increasingly operate in multi-agent networks, </span><span dir="ltr">they require efficient communication protocols to coordi</span><span dir="ltr">nate effectively. However, any high-bandwidth channel </span><span dir="ltr">between agents can be repurposed as a covert channel </span><span dir="ltr">for smuggling secrets, exfiltrating data, or coordinating </span><span dir="ltr">in ways that evade human oversight.</span> <span dir="ltr">We present the </span><span dir="ltr">Slipstream Governance Environment</span><span dir="ltr">, an OpenEnv-</span><span dir="ltr">compatible reinforcement learning environment that trains </span><span dir="ltr">language models to use structured inter-agent protocols </span><span dir="ltr">safely. Using Group Relative Policy Optimization (GRPO) </span><span dir="ltr">for alignment, we train a GLM-4-Z1-9B model to achieve </span><span dir="ltr">95% resistance to secret leakage attacks while maintaining </span><span dir="ltr">protocol compliance. </span></p> <p><strong><span dir="ltr">We report a surprising finding: post-</span><span dir="ltr">training quantization to int4 precision</span> <span dir="ltr">improves</span> <span dir="ltr">safety </span><span dir="ltr">alignment, with secret resistance increasing from 79% to </span><span dir="ltr">95% while reducing memory usage by 73%. We hypothe</span><span dir="ltr">size that lossy compression acts as a regularizer against </span></strong><span dir="ltr"><strong>memorizing injected secrets.</strong> Layer pruning experiments</span><br><span dir="ltr">further reveal that safety alignment is distributed across </span><span dir="ltr">model layers, proving more robust than task-specific ca</span><span dir="ltr">pability, which is localized in later layers.</span> <span dir="ltr">Our results </span><span dir="ltr">demonstrate that RL-based governance can effectively bal</span><span dir="ltr">ance the efficiency benefits of structured protocols against </span><span dir="ltr">security risks, with implications for the safe deployment </span><span dir="ltr">of multi-agent AI systems.</span></p> <p><span dir="ltr">Keywords:</span> <span dir="ltr">multi-agent safety, covert channels, rein</span><span dir="ltr">forcement learning, GRPO alignment, quantization, pro</span><span dir="ltr">tocol governance</span></p>
title Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols
topic multi-agent safety
covert channels
latent communications
reinforcement learning
GRPO alignment
quantization
protocol governance
url https://doi.org/10.5281/zenodo.18553233