Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Su, Zelal, Mustafaoglu, Lee, Sungyoung, Balachandar, Eshan, Miikkulainen, Risto, Pingali, Keshav
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2603.12596
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914390471606272
author Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
author_facet Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
contents Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, creating path-dependent noise. To understand this drift, we can use Fisher information geometry to decompose policy updates into signal (the natural gradient projection) and waste (the Fisher-orthogonal residual that consumes trust region budget without first-order surrogate improvement). Empirically, signal saturates but waste grows with additional epochs, creating an optimization-depth dilemma. We propose Consensus Aggregation for Policy Optimization (CAPO), which redirects compute from depth to width: $K$ PPO replicates are optimized on the same batch, differing only in minibatch shuffling order, and then aggregated into a consensus. We study aggregation in two spaces: Euclidean parameter space, and the natural parameter space of the policy distribution via the logarithmic opinion pool. In natural parameter space, the consensus provably achieves higher KL-penalized surrogate and tighter trust region compliance than the mean expert; parameter averaging inherits these guarantees approximately. On continuous control tasks, CAPO outperforms PPO and compute-matched deeper baselines under fixed sample budgets by up to 8.6x. CAPO demonstrates that policy optimization can be improved by optimizing wider, rather than deeper, without additional environment interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12596
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
Su, Zelal
Mustafaoglu
Lee, Sungyoung
Balachandar, Eshan
Miikkulainen, Risto
Pingali, Keshav
Machine Learning
Artificial Intelligence
Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, creating path-dependent noise. To understand this drift, we can use Fisher information geometry to decompose policy updates into signal (the natural gradient projection) and waste (the Fisher-orthogonal residual that consumes trust region budget without first-order surrogate improvement). Empirically, signal saturates but waste grows with additional epochs, creating an optimization-depth dilemma. We propose Consensus Aggregation for Policy Optimization (CAPO), which redirects compute from depth to width: $K$ PPO replicates are optimized on the same batch, differing only in minibatch shuffling order, and then aggregated into a consensus. We study aggregation in two spaces: Euclidean parameter space, and the natural parameter space of the policy distribution via the logarithmic opinion pool. In natural parameter space, the consensus provably achieves higher KL-penalized surrogate and tighter trust region compliance than the mean expert; parameter averaging inherits these guarantees approximately. On continuous control tasks, CAPO outperforms PPO and compute-matched deeper baselines under fixed sample budgets by up to 8.6x. CAPO demonstrates that policy optimization can be improved by optimizing wider, rather than deeper, without additional environment interactions.
title Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.12596