Unsupervised Pretraining for Fact Verification by Language Model Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bazaga, Adrián, Liò, Pietro, Micklem, Gos
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929266675941376
author Bazaga, Adrián
Liò, Pietro
Micklem, Gos
author_facet Bazaga, Adrián
Liò, Pietro
Micklem, Gos
contents Fact verification aims to verify a claim using evidence from a trustworthy knowledge base. To address this challenge, algorithms must produce features for every claim that are both semantically meaningful, and compact enough to find a semantic alignment with the source information. In contrast to previous work, which tackled the alignment problem by learning over annotated corpora of claims and their corresponding labels, we propose SFAVEL (Self-supervised Fact Verification via Language Model Distillation), a novel unsupervised pretraining framework that leverages pre-trained language models to distil self-supervised features into high-quality claim-fact alignments without the need for annotations. This is enabled by a novel contrastive loss function that encourages features to attain high-quality claim and evidence alignments whilst preserving the semantic relationships across the corpora. Notably, we present results that achieve a new state-of-the-art on FB15k-237 (+5.3% Hits@1) and FEVER (+8% accuracy) with linear evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2309_16540
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Unsupervised Pretraining for Fact Verification by Language Model Distillation
Bazaga, Adrián
Liò, Pietro
Micklem, Gos
Computation and Language
Machine Learning
Fact verification aims to verify a claim using evidence from a trustworthy knowledge base. To address this challenge, algorithms must produce features for every claim that are both semantically meaningful, and compact enough to find a semantic alignment with the source information. In contrast to previous work, which tackled the alignment problem by learning over annotated corpora of claims and their corresponding labels, we propose SFAVEL (Self-supervised Fact Verification via Language Model Distillation), a novel unsupervised pretraining framework that leverages pre-trained language models to distil self-supervised features into high-quality claim-fact alignments without the need for annotations. This is enabled by a novel contrastive loss function that encourages features to attain high-quality claim and evidence alignments whilst preserving the semantic relationships across the corpora. Notably, we present results that achieve a new state-of-the-art on FB15k-237 (+5.3% Hits@1) and FEVER (+8% accuracy) with linear evaluation.
title Unsupervised Pretraining for Fact Verification by Language Model Distillation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2309.16540