MAST: Model-Agnostic Sparsified Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Demidovich, Yury, Malinovsky, Grigory, Shulgin, Egor, Richtárik, Peter
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914235216297984
author Demidovich, Yury
Malinovsky, Grigory
Shulgin, Egor
Richtárik, Peter
author_facet Demidovich, Yury
Malinovsky, Grigory
Shulgin, Egor
Richtárik, Peter
contents We introduce a novel optimization problem formulation that departs from the conventional way of minimizing machine learning model loss as a black-box function. Unlike traditional formulations, the proposed approach explicitly incorporates an initially pre-trained model and random sketch operators, allowing for sparsification of both the model and gradient during training. We establish the insightful properties of the proposed objective function and highlight its connections to the standard formulation. Furthermore, we present several variants of the Stochastic Gradient Descent (SGD) method adapted to the new problem formulation, including SGD with general sampling, a distributed version, and SGD with variance reduction techniques. We achieve tighter convergence rates and relax assumptions, bridging the gap between theoretical principles and practical applications, covering several important techniques such as Dropout and Sparse training. This work presents promising opportunities to enhance the theoretical understanding of model training through a sparsification-aware optimization approach.
format Preprint
id arxiv_https___arxiv_org_abs_2311_16086
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MAST: Model-Agnostic Sparsified Training
Demidovich, Yury
Malinovsky, Grigory
Shulgin, Egor
Richtárik, Peter
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Optimization and Control
We introduce a novel optimization problem formulation that departs from the conventional way of minimizing machine learning model loss as a black-box function. Unlike traditional formulations, the proposed approach explicitly incorporates an initially pre-trained model and random sketch operators, allowing for sparsification of both the model and gradient during training. We establish the insightful properties of the proposed objective function and highlight its connections to the standard formulation. Furthermore, we present several variants of the Stochastic Gradient Descent (SGD) method adapted to the new problem formulation, including SGD with general sampling, a distributed version, and SGD with variance reduction techniques. We achieve tighter convergence rates and relax assumptions, bridging the gap between theoretical principles and practical applications, covering several important techniques such as Dropout and Sparse training. This work presents promising opportunities to enhance the theoretical understanding of model training through a sparsification-aware optimization approach.
title MAST: Model-Agnostic Sparsified Training
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Optimization and Control
url https://arxiv.org/abs/2311.16086