A Latent-Variable Model for Intrinsic Probing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stańczak, Karolina, Hennigen, Lucas Torroba, Williams, Adina, Cotterell, Ryan, Augenstein, Isabelle
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908481109360640
author Stańczak, Karolina
Hennigen, Lucas Torroba
Williams, Adina
Cotterell, Ryan
Augenstein, Isabelle
author_facet Stańczak, Karolina
Hennigen, Lucas Torroba
Williams, Adina
Cotterell, Ryan
Augenstein, Isabelle
contents The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that these pre-trained representations do encode some level of linguistic knowledge as they have brought about large empirical improvements on a wide variety of NLP tasks, which suggests they are learning true linguistic generalization. In this work, we focus on intrinsic probing, an analysis technique where the goal is not only to identify whether a representation encodes a linguistic attribute but also to pinpoint where this attribute is encoded. We propose a novel latent-variable formulation for constructing intrinsic probes and derive a tractable variational approximation to the log-likelihood. Our results show that our model is versatile and yields tighter mutual information estimates than two intrinsic probes previously proposed in the literature. Finally, we find empirical evidence that pre-trained representations develop a cross-lingually entangled notion of morphosyntax.
format Preprint
id arxiv_https___arxiv_org_abs_2201_08214
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle A Latent-Variable Model for Intrinsic Probing
Stańczak, Karolina
Hennigen, Lucas Torroba
Williams, Adina
Cotterell, Ryan
Augenstein, Isabelle
Computation and Language
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that these pre-trained representations do encode some level of linguistic knowledge as they have brought about large empirical improvements on a wide variety of NLP tasks, which suggests they are learning true linguistic generalization. In this work, we focus on intrinsic probing, an analysis technique where the goal is not only to identify whether a representation encodes a linguistic attribute but also to pinpoint where this attribute is encoded. We propose a novel latent-variable formulation for constructing intrinsic probes and derive a tractable variational approximation to the log-likelihood. Our results show that our model is versatile and yields tighter mutual information estimates than two intrinsic probes previously proposed in the literature. Finally, we find empirical evidence that pre-trained representations develop a cross-lingually entangled notion of morphosyntax.
title A Latent-Variable Model for Intrinsic Probing
topic Computation and Language
url https://arxiv.org/abs/2201.08214