Value Residual Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Zhanchao, Wu, Tianyi, Jiang, Zhiyun, Obeid, Fares, Lan, Zhenzhong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910993562468352
author Zhou, Zhanchao
Wu, Tianyi
Jiang, Zhiyun
Obeid, Fares
Lan, Zhenzhong
author_facet Zhou, Zhanchao
Wu, Tianyi
Jiang, Zhiyun
Obeid, Fares
Lan, Zhenzhong
contents While Transformer models have achieved remarkable success in various domains, the effectiveness of information propagation through deep networks remains a critical challenge. Standard hidden state residuals often fail to adequately preserve initial token-level information in deeper layers. This paper introduces ResFormer, a novel architecture that enhances information flow by incorporating value residual connections in addition to hidden state residuals. And a variant is SVFormer, where all layers share the first layer's value embedding. Comprehensive empirical evidence demonstrates ResFormer achieves equivalent validation loss with 16.11\% fewer model parameters and 20.3\% less training data compared to Transformer, while maintaining similar memory usage and computational cost. Besides, SVFormer reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KV-efficient methods, yielding further reductions in KV cache, with performance influenced by sequence length and cumulative learning rate.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17897
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Value Residual Learning
Zhou, Zhanchao
Wu, Tianyi
Jiang, Zhiyun
Obeid, Fares
Lan, Zhenzhong
Computation and Language
While Transformer models have achieved remarkable success in various domains, the effectiveness of information propagation through deep networks remains a critical challenge. Standard hidden state residuals often fail to adequately preserve initial token-level information in deeper layers. This paper introduces ResFormer, a novel architecture that enhances information flow by incorporating value residual connections in addition to hidden state residuals. And a variant is SVFormer, where all layers share the first layer's value embedding. Comprehensive empirical evidence demonstrates ResFormer achieves equivalent validation loss with 16.11\% fewer model parameters and 20.3\% less training data compared to Transformer, while maintaining similar memory usage and computational cost. Besides, SVFormer reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KV-efficient methods, yielding further reductions in KV cache, with performance influenced by sequence length and cumulative learning rate.
title Value Residual Learning
topic Computation and Language
url https://arxiv.org/abs/2410.17897