Your Transformer is Secretly Linear

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Razzhigaev, Anton, Mikhalchuk, Matvey, Goncharova, Elizaveta, Gerasimenko, Nikolai, Oseledets, Ivan, Dimitrov, Denis, Kuznetsov, Andrey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917671923089408
author Razzhigaev, Anton
Mikhalchuk, Matvey
Goncharova, Elizaveta
Gerasimenko, Nikolai
Oseledets, Ivan
Dimitrov, Denis
Kuznetsov, Andrey
author_facet Razzhigaev, Anton
Mikhalchuk, Matvey
Goncharova, Elizaveta
Gerasimenko, Nikolai
Oseledets, Ivan
Dimitrov, Denis
Kuznetsov, Andrey
contents This paper reveals a novel linear characteristic exclusive to transformer decoders, including models such as GPT, LLaMA, OPT, BLOOM and others. We analyze embedding transformations between sequential layers, uncovering a near-perfect linear relationship (Procrustes similarity score of 0.99). However, linearity decreases when the residual component is removed due to a consistently low output norm of the transformer layer. Our experiments show that removing or linearly approximating some of the most linear blocks of transformers does not affect significantly the loss or model performance. Moreover, in our pretraining experiments on smaller models we introduce a cosine-similarity-based regularization, aimed at reducing layer linearity. This regularization improves performance metrics on benchmarks like Tiny Stories and SuperGLUE and as well successfully decreases the linearity of the models. This study challenges the existing understanding of transformer architectures, suggesting that their operation may be more linear than previously assumed.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Your Transformer is Secretly Linear
Razzhigaev, Anton
Mikhalchuk, Matvey
Goncharova, Elizaveta
Gerasimenko, Nikolai
Oseledets, Ivan
Dimitrov, Denis
Kuznetsov, Andrey
Machine Learning
Artificial Intelligence
Computation and Language
This paper reveals a novel linear characteristic exclusive to transformer decoders, including models such as GPT, LLaMA, OPT, BLOOM and others. We analyze embedding transformations between sequential layers, uncovering a near-perfect linear relationship (Procrustes similarity score of 0.99). However, linearity decreases when the residual component is removed due to a consistently low output norm of the transformer layer. Our experiments show that removing or linearly approximating some of the most linear blocks of transformers does not affect significantly the loss or model performance. Moreover, in our pretraining experiments on smaller models we introduce a cosine-similarity-based regularization, aimed at reducing layer linearity. This regularization improves performance metrics on benchmarks like Tiny Stories and SuperGLUE and as well successfully decreases the linearity of the models. This study challenges the existing understanding of transformer architectures, suggesting that their operation may be more linear than previously assumed.
title Your Transformer is Secretly Linear
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.12250