Saved in:
Bibliographic Details
Main Authors: Su, Zelal, Mustafaoglu, Lee, Sungyoung, Balachandar, Eshan, Miikkulainen, Risto, Pingali, Keshav
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.12596
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914390471606272
author Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
author_facet Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
contents Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, creating path-dependent noise. To understand this drift, we can use Fisher information geometry to decompose policy updates into signal (the natural gradient projection) and waste (the Fisher-orthogonal residual that consumes trust region budget without first-order surrogate improvement). Empirically, signal saturates but waste grows with additional epochs, creating an optimization-depth dilemma. We propose Consensus Aggregation for Policy Optimization (CAPO), which redirects compute from depth to width: $K$ PPO replicates are optimized on the same batch, differing only in minibatch shuffling order, and then aggregated into a consensus. We study aggregation in two spaces: Euclidean parameter space, and the natural parameter space of the policy distribution via the logarithmic opinion pool. In natural parameter space, the consensus provably achieves higher KL-penalized surrogate and tighter trust region compliance than the mean expert; parameter averaging inherits these guarantees approximately. On continuous control tasks, CAPO outperforms PPO and compute-matched deeper baselines under fixed sample budgets by up to 8.6x. CAPO demonstrates that policy optimization can be improved by optimizing wider, rather than deeper, without additional environment interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12596
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
Machine Learning
Artificial Intelligence
Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, creating path-dependent noise. To understand this drift, we can use Fisher information geometry to decompose policy updates into signal (the natural gradient projection) and waste (the Fisher-orthogonal residual that consumes trust region budget without first-order surrogate improvement). Empirically, signal saturates but waste grows with additional epochs, creating an optimization-depth dilemma. We propose Consensus Aggregation for Policy Optimization (CAPO), which redirects compute from depth to width: $K$ PPO replicates are optimized on the same batch, differing only in minibatch shuffling order, and then aggregated into a consensus. We study aggregation in two spaces: Euclidean parameter space, and the natural parameter space of the policy distribution via the logarithmic opinion pool. In natural parameter space, the consensus provably achieves higher KL-penalized surrogate and tighter trust region compliance than the mean expert; parameter averaging inherits these guarantees approximately. On continuous control tasks, CAPO outperforms PPO and compute-matched deeper baselines under fixed sample budgets by up to 8.6x. CAPO demonstrates that policy optimization can be improved by optimizing wider, rather than deeper, without additional environment interactions.
title Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.12596