A multilevel approach to accelerate the training of Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lauga, Guillaume, Chaumette, Maël, Desainte-Maréville, Edgar, Lasalle, Étienne, Lebeurrier, Arthur
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916707822469120
author Lauga, Guillaume
Chaumette, Maël
Desainte-Maréville, Edgar
Lasalle, Étienne
Lebeurrier, Arthur
author_facet Lauga, Guillaume
Chaumette, Maël
Desainte-Maréville, Edgar
Lasalle, Étienne
Lebeurrier, Arthur
contents In this article, we investigate the potential of multilevel approaches to accelerate the training of transformer architectures. Using an ordinary differential equation (ODE) interpretation of these architectures, we propose an appropriate way of varying the discretization of these ODE Transformers in order to accelerate the training. We validate our approach experimentally by a comparison with the standard training procedure.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A multilevel approach to accelerate the training of Transformers
Lauga, Guillaume
Chaumette, Maël
Desainte-Maréville, Edgar
Lasalle, Étienne
Lebeurrier, Arthur
Machine Learning
Artificial Intelligence
Optimization and Control
In this article, we investigate the potential of multilevel approaches to accelerate the training of transformer architectures. Using an ordinary differential equation (ODE) interpretation of these architectures, we propose an appropriate way of varying the discretization of these ODE Transformers in order to accelerate the training. We validate our approach experimentally by a comparison with the standard training procedure.
title A multilevel approach to accelerate the training of Transformers
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2504.18590