Mixture of Experts in Image Classification: What's the Sweet Spot?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Videau, Mathurin, Leite, Alessandro, Schoenauer, Marc, Teytaud, Olivier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912669015998464
author Videau, Mathurin
Leite, Alessandro
Schoenauer, Marc
Teytaud, Olivier
author_facet Videau, Mathurin
Leite, Alessandro
Schoenauer, Marc
Teytaud, Olivier
contents Mixture-of-Experts (MoE) models have shown promising potential for parameter-efficient scaling across domains. However, their application to image classification remains limited, often requiring billion-scale datasets to be competitive. In this work, we explore the integration of MoE layers into image classification architectures using open datasets. We conduct a systematic analysis across different MoE configurations and model scales. We find that moderate parameter activation per sample provides the best trade-off between performance and efficiency. However, as the number of activated parameters increases, the benefits of MoE diminish. Our analysis yields several practical insights for vision MoE design. First, MoE layers most effectively strengthen tiny and mid-sized models, while gains taper off for large-capacity networks and do not redefine state-of-the-art ImageNet performance. Second, a Last-2 placement heuristic offers the most robust cross-architecture choice, with Every-2 slightly better for Vision Transform (ViT), and both remaining effective as data and model scale increase. Third, larger datasets (e.g., ImageNet-21k) allow more experts, up to 16, for ConvNeXt to be utilized effectively without changing placement, as increased data reduces overfitting and promotes broader expert specialization. Finally, a simple linear router performs best, suggesting that additional routing complexity yields no consistent benefit.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18322
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mixture of Experts in Image Classification: What's the Sweet Spot?
Videau, Mathurin
Leite, Alessandro
Schoenauer, Marc
Teytaud, Olivier
Computer Vision and Pattern Recognition
Machine Learning
Mixture-of-Experts (MoE) models have shown promising potential for parameter-efficient scaling across domains. However, their application to image classification remains limited, often requiring billion-scale datasets to be competitive. In this work, we explore the integration of MoE layers into image classification architectures using open datasets. We conduct a systematic analysis across different MoE configurations and model scales. We find that moderate parameter activation per sample provides the best trade-off between performance and efficiency. However, as the number of activated parameters increases, the benefits of MoE diminish. Our analysis yields several practical insights for vision MoE design. First, MoE layers most effectively strengthen tiny and mid-sized models, while gains taper off for large-capacity networks and do not redefine state-of-the-art ImageNet performance. Second, a Last-2 placement heuristic offers the most robust cross-architecture choice, with Every-2 slightly better for Vision Transform (ViT), and both remaining effective as data and model scale increase. Third, larger datasets (e.g., ImageNet-21k) allow more experts, up to 16, for ConvNeXt to be utilized effectively without changing placement, as increased data reduces overfitting and promotes broader expert specialization. Finally, a simple linear router performs best, suggesting that additional routing complexity yields no consistent benefit.
title Mixture of Experts in Image Classification: What's the Sweet Spot?
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.18322