Generalisable Agents for Neural Network Optimisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tessera, Kale-ab, Tilbury, Callum Rhys, Abramowitz, Sasha, de Kock, Ruan, Mahjoub, Omayma, Rosman, Benjamin, Hooker, Sara, Pretorius, Arnu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910378347200512
author Tessera, Kale-ab
Tilbury, Callum Rhys
Abramowitz, Sasha
de Kock, Ruan
Mahjoub, Omayma
Rosman, Benjamin
Hooker, Sara
Pretorius, Arnu
author_facet Tessera, Kale-ab
Tilbury, Callum Rhys
Abramowitz, Sasha
de Kock, Ruan
Mahjoub, Omayma
Rosman, Benjamin
Hooker, Sara
Pretorius, Arnu
contents Optimising deep neural networks is a challenging task due to complex training dynamics, high computational requirements, and long training times. To address this difficulty, we propose the framework of Generalisable Agents for Neural Network Optimisation (GANNO) -- a multi-agent reinforcement learning (MARL) approach that learns to improve neural network optimisation by dynamically and responsively scheduling hyperparameters during training. GANNO utilises an agent per layer that observes localised network dynamics and accordingly takes actions to adjust these dynamics at a layerwise level to collectively improve global performance. In this paper, we use GANNO to control the layerwise learning rate and show that the framework can yield useful and responsive schedules that are competitive with handcrafted heuristics. Furthermore, GANNO is shown to perform robustly across a wide variety of unseen initial conditions, and can successfully generalise to harder problems than it was trained on. Our work presents an overview of the opportunities that this paradigm offers for training neural networks, along with key challenges that remain to be overcome.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18598
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generalisable Agents for Neural Network Optimisation
Tessera, Kale-ab
Tilbury, Callum Rhys
Abramowitz, Sasha
de Kock, Ruan
Mahjoub, Omayma
Rosman, Benjamin
Hooker, Sara
Pretorius, Arnu
Machine Learning
Artificial Intelligence
Multiagent Systems
Optimising deep neural networks is a challenging task due to complex training dynamics, high computational requirements, and long training times. To address this difficulty, we propose the framework of Generalisable Agents for Neural Network Optimisation (GANNO) -- a multi-agent reinforcement learning (MARL) approach that learns to improve neural network optimisation by dynamically and responsively scheduling hyperparameters during training. GANNO utilises an agent per layer that observes localised network dynamics and accordingly takes actions to adjust these dynamics at a layerwise level to collectively improve global performance. In this paper, we use GANNO to control the layerwise learning rate and show that the framework can yield useful and responsive schedules that are competitive with handcrafted heuristics. Furthermore, GANNO is shown to perform robustly across a wide variety of unseen initial conditions, and can successfully generalise to harder problems than it was trained on. Our work presents an overview of the opportunities that this paradigm offers for training neural networks, along with key challenges that remain to be overcome.
title Generalisable Agents for Neural Network Optimisation
topic Machine Learning
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2311.18598