Decoupled Relative Learning Rate Schedules
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918083413671936 |
|---|---|
| author | Ludziejewski, Jan Małaśnicki, Jan Pióro, Maciej Krutul, Michał Ciebiera, Kamil Stefaniak, Maciej Krajewski, Jakub Sankowski, Piotr Cygan, Marek Adamczewski, Kamil Jaszczur, Sebastian |
| author_facet | Ludziejewski, Jan Małaśnicki, Jan Pióro, Maciej Krutul, Michał Ciebiera, Kamil Stefaniak, Maciej Krajewski, Jakub Sankowski, Piotr Cygan, Marek Adamczewski, Kamil Jaszczur, Sebastian |
| contents | In this work, we introduce a novel approach for optimizing LLM training by adjusting learning rates across weights of different components in Transformer models. Traditional methods often apply a uniform learning rate across all network layers, potentially overlooking the unique dynamics of each part. Remarkably, our introduced relative learning rates, RLRS, method accelerates the training process by up to $23\%$, particularly in complex models such as Mixture of Experts (MoE). Hyperparameters of RLRS can be efficiently tuned on smaller models and then effectively reused on models up to $27\times$ larger. This simple and effective method results in a substantial reduction in training time and computational resources, offering a practical and scalable solution for optimizing large-scale neural networks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_03526 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Decoupled Relative Learning Rate Schedules Ludziejewski, Jan Małaśnicki, Jan Pióro, Maciej Krutul, Michał Ciebiera, Kamil Stefaniak, Maciej Krajewski, Jakub Sankowski, Piotr Cygan, Marek Adamczewski, Kamil Jaszczur, Sebastian Machine Learning In this work, we introduce a novel approach for optimizing LLM training by adjusting learning rates across weights of different components in Transformer models. Traditional methods often apply a uniform learning rate across all network layers, potentially overlooking the unique dynamics of each part. Remarkably, our introduced relative learning rates, RLRS, method accelerates the training process by up to $23\%$, particularly in complex models such as Mixture of Experts (MoE). Hyperparameters of RLRS can be efficiently tuned on smaller models and then effectively reused on models up to $27\times$ larger. This simple and effective method results in a substantial reduction in training time and computational resources, offering a practical and scalable solution for optimizing large-scale neural networks. |
| title | Decoupled Relative Learning Rate Schedules |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2507.03526 |