Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Bei, Zheng, Tong, Wang, Rui, Liu, Jiahao, Guo, Qingyan, Guo, Junliang, Tan, Xu, Xiao, Tong, Zhu, Jingbo, Wang, Jingang, Cai, Xunliang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929577068068864
author Li, Bei
Zheng, Tong
Wang, Rui
Liu, Jiahao
Guo, Qingyan
Guo, Junliang
Tan, Xu
Xiao, Tong
Zhu, Jingbo
Wang, Jingang
Cai, Xunliang
author_facet Li, Bei
Zheng, Tong
Wang, Rui
Liu, Jiahao
Guo, Qingyan
Guo, Junliang
Tan, Xu
Xiao, Tong
Zhu, Jingbo
Wang, Jingang
Cai, Xunliang
contents Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution.'' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30.95 and 44.27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3.8B DeepNet by an average of 2.9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5.7 accuracy points on the LM Harness Evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03042
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
Li, Bei
Zheng, Tong
Wang, Rui
Liu, Jiahao
Guo, Qingyan
Guo, Junliang
Tan, Xu
Xiao, Tong
Zhu, Jingbo
Wang, Jingang
Cai, Xunliang
Computation and Language
Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution.'' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30.95 and 44.27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3.8B DeepNet by an average of 2.9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5.7 accuracy points on the LM Harness Evaluation.
title Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
topic Computation and Language
url https://arxiv.org/abs/2411.03042