Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols
Fuente:
Zenodo
Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866902236393635840 |
|---|---|
| author | Maio, Anthony D. |
| author_facet | Maio, Anthony D. |
| contents | <p><span dir="ltr">As AI agents increasingly operate in multi-agent networks, </span><span dir="ltr">they require efficient communication protocols to coordi</span><span dir="ltr">nate effectively. However, any high-bandwidth channel </span><span dir="ltr">between agents can be repurposed as a covert channel </span><span dir="ltr">for smuggling secrets, exfiltrating data, or coordinating </span><span dir="ltr">in ways that evade human oversight.</span> <span dir="ltr">We present the </span><span dir="ltr">Slipstream Governance Environment</span><span dir="ltr">, an OpenEnv-</span><span dir="ltr">compatible reinforcement learning environment that trains </span><span dir="ltr">language models to use structured inter-agent protocols </span><span dir="ltr">safely. Using Group Relative Policy Optimization (GRPO) </span><span dir="ltr">for alignment, we train a GLM-4-Z1-9B model to achieve </span><span dir="ltr">95% resistance to secret leakage attacks while maintaining </span><span dir="ltr">protocol compliance. </span></p> <p><strong><span dir="ltr">We report a surprising finding: post-</span><span dir="ltr">training quantization to int4 precision</span> <span dir="ltr">improves</span> <span dir="ltr">safety </span><span dir="ltr">alignment, with secret resistance increasing from 79% to </span><span dir="ltr">95% while reducing memory usage by 73%. We hypothe</span><span dir="ltr">size that lossy compression acts as a regularizer against </span></strong><span dir="ltr"><strong>memorizing injected secrets.</strong> Layer pruning experiments</span><br><span dir="ltr">further reveal that safety alignment is distributed across </span><span dir="ltr">model layers, proving more robust than task-specific ca</span><span dir="ltr">pability, which is localized in later layers.</span> <span dir="ltr">Our results </span><span dir="ltr">demonstrate that RL-based governance can effectively bal</span><span dir="ltr">ance the efficiency benefits of structured protocols against </span><span dir="ltr">security risks, with implications for the safe deployment </span><span dir="ltr">of multi-agent AI systems.</span></p> <p><span dir="ltr">Keywords:</span> <span dir="ltr">multi-agent safety, covert channels, rein</span><span dir="ltr">forcement learning, GRPO alignment, quantization, pro</span><span dir="ltr">tocol governance</span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18553233 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols Maio, Anthony D. multi-agent safety covert channels latent communications reinforcement learning GRPO alignment quantization protocol governance <p><span dir="ltr">As AI agents increasingly operate in multi-agent networks, </span><span dir="ltr">they require efficient communication protocols to coordi</span><span dir="ltr">nate effectively. However, any high-bandwidth channel </span><span dir="ltr">between agents can be repurposed as a covert channel </span><span dir="ltr">for smuggling secrets, exfiltrating data, or coordinating </span><span dir="ltr">in ways that evade human oversight.</span> <span dir="ltr">We present the </span><span dir="ltr">Slipstream Governance Environment</span><span dir="ltr">, an OpenEnv-</span><span dir="ltr">compatible reinforcement learning environment that trains </span><span dir="ltr">language models to use structured inter-agent protocols </span><span dir="ltr">safely. Using Group Relative Policy Optimization (GRPO) </span><span dir="ltr">for alignment, we train a GLM-4-Z1-9B model to achieve </span><span dir="ltr">95% resistance to secret leakage attacks while maintaining </span><span dir="ltr">protocol compliance. </span></p> <p><strong><span dir="ltr">We report a surprising finding: post-</span><span dir="ltr">training quantization to int4 precision</span> <span dir="ltr">improves</span> <span dir="ltr">safety </span><span dir="ltr">alignment, with secret resistance increasing from 79% to </span><span dir="ltr">95% while reducing memory usage by 73%. We hypothe</span><span dir="ltr">size that lossy compression acts as a regularizer against </span></strong><span dir="ltr"><strong>memorizing injected secrets.</strong> Layer pruning experiments</span><br><span dir="ltr">further reveal that safety alignment is distributed across </span><span dir="ltr">model layers, proving more robust than task-specific ca</span><span dir="ltr">pability, which is localized in later layers.</span> <span dir="ltr">Our results </span><span dir="ltr">demonstrate that RL-based governance can effectively bal</span><span dir="ltr">ance the efficiency benefits of structured protocols against </span><span dir="ltr">security risks, with implications for the safe deployment </span><span dir="ltr">of multi-agent AI systems.</span></p> <p><span dir="ltr">Keywords:</span> <span dir="ltr">multi-agent safety, covert channels, rein</span><span dir="ltr">forcement learning, GRPO alignment, quantization, pro</span><span dir="ltr">tocol governance</span></p> |
| title | Training AI Agents to Communicate Safely: Reinforcement Learning for Covert Channel Prevention in Inter-Agent Protocols |
| topic | multi-agent safety covert channels latent communications reinforcement learning GRPO alignment quantization protocol governance |
| url | https://doi.org/10.5281/zenodo.18553233 |