Reversed Attention: On The Gradient Descent Of Attention Layers In GPT

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Katz, Shahar, Wolf, Lior
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916538658848768
author Katz, Shahar
Wolf, Lior
author_facet Katz, Shahar
Wolf, Lior
contents The success of Transformer-based Language Models (LMs) stems from their attention mechanism. While this mechanism has been extensively studied in explainability research, particularly through the attention values obtained during the forward pass of LMs, the backward pass of attention has been largely overlooked. In this work, we study the mathematics of the backward pass of attention, revealing that it implicitly calculates an attention matrix we refer to as "Reversed Attention". We examine the properties of Reversed Attention and demonstrate its ability to elucidate the models' behavior and edit dynamics. In an experimental setup, we showcase the ability of Reversed Attention to directly alter the forward pass of attention, without modifying the model's weights, using a novel method called "attention patching". In addition to enhancing the comprehension of how LM configure attention layers during backpropagation, Reversed Attention maps contribute to a more interpretable backward pass.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17019
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
Katz, Shahar
Wolf, Lior
Computation and Language
The success of Transformer-based Language Models (LMs) stems from their attention mechanism. While this mechanism has been extensively studied in explainability research, particularly through the attention values obtained during the forward pass of LMs, the backward pass of attention has been largely overlooked. In this work, we study the mathematics of the backward pass of attention, revealing that it implicitly calculates an attention matrix we refer to as "Reversed Attention". We examine the properties of Reversed Attention and demonstrate its ability to elucidate the models' behavior and edit dynamics. In an experimental setup, we showcase the ability of Reversed Attention to directly alter the forward pass of attention, without modifying the model's weights, using a novel method called "attention patching". In addition to enhancing the comprehension of how LM configure attention layers during backpropagation, Reversed Attention maps contribute to a more interpretable backward pass.
title Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
topic Computation and Language
url https://arxiv.org/abs/2412.17019