Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nastase, Vivi, Merlo, Paola
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913445904908288
author Nastase, Vivi
Merlo, Paola
author_facet Nastase, Vivi
Merlo, Paola
contents Analyses of transformer-based models have shown that they encode a variety of linguistic information from their textual input. While these analyses have shed a light on the relation between linguistic information on one side, and internal architecture and parameters on the other, a question remains unanswered: how is this linguistic information reflected in sentence embeddings? Using datasets consisting of sentences with known structure, we test to what degree information about chunks (in particular noun, verb or prepositional phrases), such as grammatical number, or semantic role, can be localized in sentence embeddings. Our results show that such information is not distributed over the entire sentence embedding, but rather it is encoded in specific regions. Understanding how the information from an input text is compressed into sentence embeddings helps understand current transformer models and help build future explainable neural models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18119
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification
Nastase, Vivi
Merlo, Paola
Computation and Language
68T50
I.2.7
Analyses of transformer-based models have shown that they encode a variety of linguistic information from their textual input. While these analyses have shed a light on the relation between linguistic information on one side, and internal architecture and parameters on the other, a question remains unanswered: how is this linguistic information reflected in sentence embeddings? Using datasets consisting of sentences with known structure, we test to what degree information about chunks (in particular noun, verb or prepositional phrases), such as grammatical number, or semantic role, can be localized in sentence embeddings. Our results show that such information is not distributed over the entire sentence embedding, but rather it is encoded in specific regions. Understanding how the information from an input text is compressed into sentence embeddings helps understand current transformer models and help build future explainable neural models.
title Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2407.18119