More Experts Than Galaxies: Conditionally-overlapping Experts With Biologically-Inspired Fixed Routing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaier, Sagi, Pereira, Francisco, von der Wense, Katharina, Hunter, Lawrence E, Jones, Matt
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929710741585920
author Shaier, Sagi
Pereira, Francisco
von der Wense, Katharina
Hunter, Lawrence E
Jones, Matt
author_facet Shaier, Sagi
Pereira, Francisco
von der Wense, Katharina
Hunter, Lawrence E
Jones, Matt
contents The evolution of biological neural systems has led to both modularity and sparse coding, which enables energy efficiency and robustness across the diversity of tasks in the lifespan. In contrast, standard neural networks rely on dense, non-specialized architectures, where all model parameters are simultaneously updated to learn multiple tasks, leading to interference. Current sparse neural network approaches aim to alleviate this issue but are hindered by limitations such as 1) trainable gating functions that cause representation collapse, 2) disjoint experts that result in redundant computation and slow learning, and 3) reliance on explicit input or task IDs that limit flexibility and scalability. In this paper we propose Conditionally Overlapping Mixture of ExperTs (COMET), a general deep learning method that addresses these challenges by inducing a modular, sparse architecture with an exponential number of overlapping experts. COMET replaces the trainable gating function used in Sparse Mixture of Experts with a fixed, biologically inspired random projection applied to individual input representations. This design causes the degree of expert overlap to depend on input similarity, so that similar inputs tend to share more parameters. This results in faster learning per update step and improved out-of-sample generalization. We demonstrate the effectiveness of COMET on a range of tasks, including image classification, language modeling, and regression, using several popular deep learning architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08003
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle More Experts Than Galaxies: Conditionally-overlapping Experts With Biologically-Inspired Fixed Routing
Shaier, Sagi
Pereira, Francisco
von der Wense, Katharina
Hunter, Lawrence E
Jones, Matt
Machine Learning
The evolution of biological neural systems has led to both modularity and sparse coding, which enables energy efficiency and robustness across the diversity of tasks in the lifespan. In contrast, standard neural networks rely on dense, non-specialized architectures, where all model parameters are simultaneously updated to learn multiple tasks, leading to interference. Current sparse neural network approaches aim to alleviate this issue but are hindered by limitations such as 1) trainable gating functions that cause representation collapse, 2) disjoint experts that result in redundant computation and slow learning, and 3) reliance on explicit input or task IDs that limit flexibility and scalability. In this paper we propose Conditionally Overlapping Mixture of ExperTs (COMET), a general deep learning method that addresses these challenges by inducing a modular, sparse architecture with an exponential number of overlapping experts. COMET replaces the trainable gating function used in Sparse Mixture of Experts with a fixed, biologically inspired random projection applied to individual input representations. This design causes the degree of expert overlap to depend on input similarity, so that similar inputs tend to share more parameters. This results in faster learning per update step and improved out-of-sample generalization. We demonstrate the effectiveness of COMET on a range of tasks, including image classification, language modeling, and regression, using several popular deep learning architectures.
title More Experts Than Galaxies: Conditionally-overlapping Experts With Biologically-Inspired Fixed Routing
topic Machine Learning
url https://arxiv.org/abs/2410.08003