Decomposing Query-Key Feature Interactions Using Contrastive Covariances

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Andrew, Belinkov, Yonatan, Viégas, Fernanda, Wattenberg, Martin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912877601882112
author Lee, Andrew
Belinkov, Yonatan
Viégas, Fernanda
Wattenberg, Martin
author_facet Lee, Andrew
Belinkov, Yonatan
Viégas, Fernanda
Wattenberg, Martin
contents Despite the central role of attention heads in Transformers, we lack tools to understand why a model attends to a particular token. To address this, we study the query-key (QK) space -- the bilinear joint embedding space between queries and keys. We present a contrastive covariance method to decompose the QK space into low-rank, human-interpretable components. It is when features in keys and queries align in these low-rank subspaces that high attention scores are produced. We first study our method both analytically and empirically in a simplified setting. We then apply our method to large language models to identify human-interpretable QK subspaces for categorical semantic features and binding features. Finally, we demonstrate how attention scores can be attributed to our identified features.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04752
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Decomposing Query-Key Feature Interactions Using Contrastive Covariances
Lee, Andrew
Belinkov, Yonatan
Viégas, Fernanda
Wattenberg, Martin
Machine Learning
Despite the central role of attention heads in Transformers, we lack tools to understand why a model attends to a particular token. To address this, we study the query-key (QK) space -- the bilinear joint embedding space between queries and keys. We present a contrastive covariance method to decompose the QK space into low-rank, human-interpretable components. It is when features in keys and queries align in these low-rank subspaces that high attention scores are produced. We first study our method both analytically and empirically in a simplified setting. We then apply our method to large language models to identify human-interpretable QK subspaces for categorical semantic features and binding features. Finally, we demonstrate how attention scores can be attributed to our identified features.
title Decomposing Query-Key Feature Interactions Using Contrastive Covariances
topic Machine Learning
url https://arxiv.org/abs/2602.04752