Does equivariance matter at scale?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Brehmer, Johann, Behrends, Sönke, de Haan, Pim, Cohen, Taco
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911078809600000
author Brehmer, Johann
Behrends, Sönke
de Haan, Pim
Cohen, Taco
author_facet Brehmer, Johann
Behrends, Sönke
de Haan, Pim
Cohen, Taco
contents Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does equivariance matter at scale?
Brehmer, Johann
Behrends, Sönke
de Haan, Pim
Cohen, Taco
Machine Learning
Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.
title Does equivariance matter at scale?
topic Machine Learning
url https://arxiv.org/abs/2410.23179