Careful with that Scalpel: Improving Gradient Surgery with an EMA

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hsieh, Yu-Guan, Thornton, James, Ndiaye, Eugene, Klein, Michal, Cuturi, Marco, Ablin, Pierre
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913720697880576
author Hsieh, Yu-Guan
Thornton, James
Ndiaye, Eugene
Klein, Michal
Cuturi, Marco
Ablin, Pierre
author_facet Hsieh, Yu-Guan
Thornton, James
Ndiaye, Eugene
Klein, Michal
Cuturi, Marco
Ablin, Pierre
contents Beyond minimizing a single training loss, many deep learning estimation pipelines rely on an auxiliary objective to quantify and encourage desirable properties of the model (e.g. performance on another dataset, robustness, agreement with a prior). Although the simplest approach to incorporating an auxiliary loss is to sum it with the training loss as a regularizer, recent works have shown that one can improve performance by blending the gradients beyond a simple sum; this is known as gradient surgery. We cast the problem as a constrained minimization problem where the auxiliary objective is minimized among the set of minimizers of the training loss. To solve this bilevel problem, we follow a parameter update direction that combines the training loss gradient and the orthogonal projection of the auxiliary gradient to the training gradient. In a setting where gradients come from mini-batches, we explain how, using a moving average of the training loss gradients, we can carefully maintain this critical orthogonality property. We demonstrate that our method, Bloop, can lead to much better performances on NLP and vision experiments than other gradient surgery methods without EMA.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02998
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Careful with that Scalpel: Improving Gradient Surgery with an EMA
Hsieh, Yu-Guan
Thornton, James
Ndiaye, Eugene
Klein, Michal
Cuturi, Marco
Ablin, Pierre
Machine Learning
Beyond minimizing a single training loss, many deep learning estimation pipelines rely on an auxiliary objective to quantify and encourage desirable properties of the model (e.g. performance on another dataset, robustness, agreement with a prior). Although the simplest approach to incorporating an auxiliary loss is to sum it with the training loss as a regularizer, recent works have shown that one can improve performance by blending the gradients beyond a simple sum; this is known as gradient surgery. We cast the problem as a constrained minimization problem where the auxiliary objective is minimized among the set of minimizers of the training loss. To solve this bilevel problem, we follow a parameter update direction that combines the training loss gradient and the orthogonal projection of the auxiliary gradient to the training gradient. In a setting where gradients come from mini-batches, we explain how, using a moving average of the training loss gradients, we can carefully maintain this critical orthogonality property. We demonstrate that our method, Bloop, can lead to much better performances on NLP and vision experiments than other gradient surgery methods without EMA.
title Careful with that Scalpel: Improving Gradient Surgery with an EMA
topic Machine Learning
url https://arxiv.org/abs/2402.02998