An Analysis of Multilingual FActScore

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vu, Kim Trong, Krumdick, Michael, Reddy, Varshini, Dernoncourt, Franck, Lai, Viet Dac
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909232820912128
author Vu, Kim Trong
Krumdick, Michael
Reddy, Varshini
Dernoncourt, Franck
Lai, Viet Dac
author_facet Vu, Kim Trong
Krumdick, Michael
Reddy, Varshini
Dernoncourt, Franck
Lai, Viet Dac
contents FActScore has gained popularity as a metric to estimate the factuality of long-form texts generated by Large Language Models (LLMs) in English. However, there has not been any work in studying the behavior of FActScore in other languages. This paper studies the limitations of each component in the four-component pipeline of FActScore in the multilingual setting. We introduce a new dataset for FActScore on texts generated by strong multilingual LLMs. Our evaluation shows that LLMs exhibit distinct behaviors in both fact extraction and fact scoring tasks. No LLM produces consistent and reliable FActScore across languages with varying levels of resources. We also find that the knowledge source plays an important role in the quality of the estimated FActScore. Using Wikipedia as the knowledge source may hinder the true FActScore of long-form text due to its limited coverage in medium- and low-resource languages. We also incorporate three mitigations to our knowledge source that ultimately improve FActScore estimation across all languages.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Analysis of Multilingual FActScore
Vu, Kim Trong
Krumdick, Michael
Reddy, Varshini
Dernoncourt, Franck
Lai, Viet Dac
Computation and Language
FActScore has gained popularity as a metric to estimate the factuality of long-form texts generated by Large Language Models (LLMs) in English. However, there has not been any work in studying the behavior of FActScore in other languages. This paper studies the limitations of each component in the four-component pipeline of FActScore in the multilingual setting. We introduce a new dataset for FActScore on texts generated by strong multilingual LLMs. Our evaluation shows that LLMs exhibit distinct behaviors in both fact extraction and fact scoring tasks. No LLM produces consistent and reliable FActScore across languages with varying levels of resources. We also find that the knowledge source plays an important role in the quality of the estimated FActScore. Using Wikipedia as the knowledge source may hinder the true FActScore of long-form text due to its limited coverage in medium- and low-resource languages. We also incorporate three mitigations to our knowledge source that ultimately improve FActScore estimation across all languages.
title An Analysis of Multilingual FActScore
topic Computation and Language
url https://arxiv.org/abs/2406.19415