Adaptive Ensemble Aggregation for Actor-Critics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Werge, Nicklas, Wu, Yi-Shan, Haussmann, Manuel, Tasdighi, Bahareh, Kandemir, Melih
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914532979376128
author Werge, Nicklas
Wu, Yi-Shan
Haussmann, Manuel
Tasdighi, Bahareh
Kandemir, Melih
author_facet Werge, Nicklas
Wu, Yi-Shan
Haussmann, Manuel
Tasdighi, Bahareh
Kandemir, Melih
contents Ensembles are ubiquitous in off-policy actor-critic learning, yet their efficacy depends critically on how they are aggregated. Current methods typically rely on static rules or task-specific hyperparameters to balance overestimation bias and variance, leaving the challenge of a truly adaptive approach open. We introduce Adaptive Ensemble Aggregation (AEA), an algorithm that dynamically constructs ensemble-based targets for both critic and actor updates directly from training dynamics. We prove that AEA converges to a unique equilibrium where the aggregation parameter minimizes value estimation error within a defined stability region. Theoretically, we establish that AEA achieves a shrinkage property where the estimation bias vanishes as the total ensemble size grows. Unlike subset-based methods like REDQ, which hit an information bottleneck determined by a fixed variance floor regardless of the ensemble size, AEA exploits the full ensemble to achieve optimal variance reduction-scaling inversely with the total number of models-and maximal Fisher information. Furthermore, we provide a formal guarantee for monotonic policy improvement under this adaptive regime. Extensive evaluations on various continuous control tasks demonstrate that AEA outperforms, on the majority of tasks, state-of-the-art baselines, providing a robust and self-calibrating framework for ensemble-based reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Ensemble Aggregation for Actor-Critics
Werge, Nicklas
Wu, Yi-Shan
Haussmann, Manuel
Tasdighi, Bahareh
Kandemir, Melih
Machine Learning
Ensembles are ubiquitous in off-policy actor-critic learning, yet their efficacy depends critically on how they are aggregated. Current methods typically rely on static rules or task-specific hyperparameters to balance overestimation bias and variance, leaving the challenge of a truly adaptive approach open. We introduce Adaptive Ensemble Aggregation (AEA), an algorithm that dynamically constructs ensemble-based targets for both critic and actor updates directly from training dynamics. We prove that AEA converges to a unique equilibrium where the aggregation parameter minimizes value estimation error within a defined stability region. Theoretically, we establish that AEA achieves a shrinkage property where the estimation bias vanishes as the total ensemble size grows. Unlike subset-based methods like REDQ, which hit an information bottleneck determined by a fixed variance floor regardless of the ensemble size, AEA exploits the full ensemble to achieve optimal variance reduction-scaling inversely with the total number of models-and maximal Fisher information. Furthermore, we provide a formal guarantee for monotonic policy improvement under this adaptive regime. Extensive evaluations on various continuous control tasks demonstrate that AEA outperforms, on the majority of tasks, state-of-the-art baselines, providing a robust and self-calibrating framework for ensemble-based reinforcement learning.
title Adaptive Ensemble Aggregation for Actor-Critics
topic Machine Learning
url https://arxiv.org/abs/2507.23501