Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koran, Eugene, Yun, Yejun, Tetef, Samantha, Arnav, Benjamin, Bernabeu-Pérez, Pablo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917507423535104
author Koran, Eugene
Yun, Yejun
Tetef, Samantha
Arnav, Benjamin
Bernabeu-Pérez, Pablo
author_facet Koran, Eugene
Yun, Yejun
Tetef, Samantha
Arnav, Benjamin
Bernabeu-Pérez, Pablo
contents As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned with user intent. Monitoring agent actions is a key safety mechanism, yet reliable monitors remain difficult to build and the scale of these systems makes human oversight impractical. We show that combining signals from diverse monitors into an ensemble improves detection of misaligned actions. We build 12 GPT-4.1-Mini monitors using both prompting and fine-tuning strategies. We evaluate them on coding tasks where candidate solutions pass standard tests but fail on adversarial inputs. In this setting, diverse ensembles outperform both individual monitors and homogeneous ensembles. Our best 3-monitor ensemble achieves 2.4x greater detection performance gain compared to an ensemble composed of three identical monitors, with the same ensemble performing strongly on an independent dataset. We contend that these results show that diversity - not scale - drives gains. The best ensembles combine strong individual performance with low correlation between monitors. Furthermore, fine-tuned monitors appear in every top-performing ensemble and maintain this advantage on out-of-distribution attack types, suggesting that fine-tuning enables detection capabilities that prompting alone does not elicit. These results support ensemble monitoring as a practical AI control strategy for safety gains at reasonable inference costs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
Koran, Eugene
Yun, Yejun
Tetef, Samantha
Arnav, Benjamin
Bernabeu-Pérez, Pablo
Artificial Intelligence
As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned with user intent. Monitoring agent actions is a key safety mechanism, yet reliable monitors remain difficult to build and the scale of these systems makes human oversight impractical. We show that combining signals from diverse monitors into an ensemble improves detection of misaligned actions. We build 12 GPT-4.1-Mini monitors using both prompting and fine-tuning strategies. We evaluate them on coding tasks where candidate solutions pass standard tests but fail on adversarial inputs. In this setting, diverse ensembles outperform both individual monitors and homogeneous ensembles. Our best 3-monitor ensemble achieves 2.4x greater detection performance gain compared to an ensemble composed of three identical monitors, with the same ensemble performing strongly on an independent dataset. We contend that these results show that diversity - not scale - drives gains. The best ensembles combine strong individual performance with low correlation between monitors. Furthermore, fine-tuned monitors appear in every top-performing ensemble and maintain this advantage on out-of-distribution attack types, suggesting that fine-tuning enables detection capabilities that prompting alone does not elicit. These results support ensemble monitoring as a practical AI control strategy for safety gains at reasonable inference costs.
title Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
topic Artificial Intelligence
url https://arxiv.org/abs/2605.15377