Looking Beyond The Top-1: Transformers Determine Top Tokens In Order

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lioubashevski, Daria, Schlank, Tomer, Stanovsky, Gabriel, Goldstein, Ariel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909366556295168
author Lioubashevski, Daria
Schlank, Tomer
Stanovsky, Gabriel
Goldstein, Ariel
author_facet Lioubashevski, Daria
Schlank, Tomer
Stanovsky, Gabriel
Goldstein, Ariel
contents Understanding the inner workings of Transformers is crucial for achieving more accurate and efficient predictions. In this work, we analyze the computation performed by Transformers in the layers after the top-1 prediction has become fixed, which has been previously referred to as the "saturation event". We expand the concept of saturation events for top-k tokens, demonstrating that similar saturation events occur across language, vision, and speech models. We find that these saturation events happen in order of the corresponding tokens' ranking, i.e., the model first decides on the top ranking token, then the second highest ranking token, and so on. This phenomenon seems intrinsic to the Transformer architecture, occurring across different architectural variants (decoder-only, encoder-only, and to a lesser extent full-Transformer), and even in untrained Transformers. We propose an underlying mechanism of task transition for this sequential saturation, where task k corresponds to predicting the k-th most probable token, and the saturation events are in fact discrete transitions between the tasks. In support of this we show that it is possible to predict the current task from hidden layer embedding. Furthermore, using an intervention method we demonstrate that we can cause the model to switch from one task to the next. Finally, leveraging our findings, we introduce a novel token-level early-exit strategy, which surpasses existing methods in balancing performance and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20210
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
Lioubashevski, Daria
Schlank, Tomer
Stanovsky, Gabriel
Goldstein, Ariel
Computation and Language
Machine Learning
Understanding the inner workings of Transformers is crucial for achieving more accurate and efficient predictions. In this work, we analyze the computation performed by Transformers in the layers after the top-1 prediction has become fixed, which has been previously referred to as the "saturation event". We expand the concept of saturation events for top-k tokens, demonstrating that similar saturation events occur across language, vision, and speech models. We find that these saturation events happen in order of the corresponding tokens' ranking, i.e., the model first decides on the top ranking token, then the second highest ranking token, and so on. This phenomenon seems intrinsic to the Transformer architecture, occurring across different architectural variants (decoder-only, encoder-only, and to a lesser extent full-Transformer), and even in untrained Transformers. We propose an underlying mechanism of task transition for this sequential saturation, where task k corresponds to predicting the k-th most probable token, and the saturation events are in fact discrete transitions between the tasks. In support of this we show that it is possible to predict the current task from hidden layer embedding. Furthermore, using an intervention method we demonstrate that we can cause the model to switch from one task to the next. Finally, leveraging our findings, we introduce a novel token-level early-exit strategy, which surpasses existing methods in balancing performance and efficiency.
title Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.20210