Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913526675668992 |
|---|---|
| author | Li, Qingyang Ke, Weimao |
| author_facet | Li, Qingyang Ke, Weimao |
| contents | This paper examines the pivotal role of dropout techniques in mitigating overfitting in language model training. It conducts a comprehensive investigation into the influence of variable dropout rates on both individual layers and residual connections within the context of language modeling. Our study conducts training of a decoder implementation on the classic Tiny Shakespeare data to examine the effects of the adjustments on training efficiency and validation error. Results not only confirm the benefits of dropout for regularization and residuals for convergence, but also reveal their interesting interactions. There exists an important trade-off between the depth of residual connections and the dropout on these connections for optimal deep neural network convergence and generalization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_01019 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training Li, Qingyang Ke, Weimao Computation and Language Machine Learning This paper examines the pivotal role of dropout techniques in mitigating overfitting in language model training. It conducts a comprehensive investigation into the influence of variable dropout rates on both individual layers and residual connections within the context of language modeling. Our study conducts training of a decoder implementation on the classic Tiny Shakespeare data to examine the effects of the adjustments on training efficiency and validation error. Results not only confirm the benefits of dropout for regularization and residuals for convergence, but also reveal their interesting interactions. There exists an important trade-off between the depth of residual connections and the dropout on these connections for optimal deep neural network convergence and generalization. |
| title | Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2410.01019 |