The Structure of Relation Decoding Linear Operators in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Christ, Miranda Anna, Csiszárik, Adrián, Becsó, Gergely, Varga, Dániel
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917050445725696
author Christ, Miranda Anna
Csiszárik, Adrián
Becsó, Gergely
Varga, Dániel
author_facet Christ, Miranda Anna
Csiszárik, Adrián
Becsó, Gergely
Varga, Dániel
contents This paper investigates the structure of linear operators introduced in Hernandez et al. [2023] that decode specific relational facts in transformer language models. We extend their single-relation findings to a collection of relations and systematically chart their organization. We show that such collections of relation decoders can be highly compressed by simple order-3 tensor networks without significant loss in decoding accuracy. To explain this surprising redundancy, we develop a cross-evaluation protocol, in which we apply each linear decoder operator to the subjects of every other relation. Our results reveal that these linear maps do not encode distinct relations, but extract recurring, coarse-grained semantic properties (e.g., country of capital city and country of food are both in the country-of-X property). This property-centric structure clarifies both the operators' compressibility and highlights why they generalize only to new relations that are semantically close. Our findings thus interpret linear relational decoding in transformer language models as primarily property-based, rather than relation-specific.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Structure of Relation Decoding Linear Operators in Large Language Models
Christ, Miranda Anna
Csiszárik, Adrián
Becsó, Gergely
Varga, Dániel
Computation and Language
Artificial Intelligence
Machine Learning
This paper investigates the structure of linear operators introduced in Hernandez et al. [2023] that decode specific relational facts in transformer language models. We extend their single-relation findings to a collection of relations and systematically chart their organization. We show that such collections of relation decoders can be highly compressed by simple order-3 tensor networks without significant loss in decoding accuracy. To explain this surprising redundancy, we develop a cross-evaluation protocol, in which we apply each linear decoder operator to the subjects of every other relation. Our results reveal that these linear maps do not encode distinct relations, but extract recurring, coarse-grained semantic properties (e.g., country of capital city and country of food are both in the country-of-X property). This property-centric structure clarifies both the operators' compressibility and highlights why they generalize only to new relations that are semantically close. Our findings thus interpret linear relational decoding in transformer language models as primarily property-based, rather than relation-specific.
title The Structure of Relation Decoding Linear Operators in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.26543