OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Iqbal, Hasan, Wang, Yuxia, Wang, Minghan, Georgiev, Georgi, Geng, Jiahui, Gurevych, Iryna, Nakov, Preslav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911237718147072
author Iqbal, Hasan
Wang, Yuxia
Wang, Minghan
Georgiev, Georgi
Geng, Jiahui
Gurevych, Iryna
Nakov, Preslav
author_facet Iqbal, Hasan
Wang, Yuxia
Wang, Minghan
Georgiev, Georgi
Geng, Jiahui
Gurevych, Iryna
Nakov, Preslav
contents The increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate. This is difficult as it requires assessing the factuality of free-form open-domain responses. While there has been a lot of research on this topic, different papers use different evaluation benchmarks and measures, which makes them hard to compare and hampers future progress. To mitigate these issues, we developed OpenFactCheck, a unified framework, with three modules: (i) RESPONSEEVAL, which allows users to easily customize an automatic fact-checking system and to assess the factuality of all claims in an input document using that system, (ii) LLMEVAL, which assesses the overall factuality of an LLM, and (iii) CHECKEREVAL, a module to evaluate automatic fact-checking systems. OpenFactCheck is open-sourced (https://github.com/mbzuai-nlp/openfactcheck) and publicly released as a Python library (https://pypi.org/project/openfactcheck/) and also as a web service (http://app.openfactcheck.com). A video describing the system is available at https://youtu.be/-i9VKL0HleI.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11832
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
Iqbal, Hasan
Wang, Yuxia
Wang, Minghan
Georgiev, Georgi
Geng, Jiahui
Gurevych, Iryna
Nakov, Preslav
Computation and Language
Artificial Intelligence
I.2.7
The increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate. This is difficult as it requires assessing the factuality of free-form open-domain responses. While there has been a lot of research on this topic, different papers use different evaluation benchmarks and measures, which makes them hard to compare and hampers future progress. To mitigate these issues, we developed OpenFactCheck, a unified framework, with three modules: (i) RESPONSEEVAL, which allows users to easily customize an automatic fact-checking system and to assess the factuality of all claims in an input document using that system, (ii) LLMEVAL, which assesses the overall factuality of an LLM, and (iii) CHECKEREVAL, a module to evaluate automatic fact-checking systems. OpenFactCheck is open-sourced (https://github.com/mbzuai-nlp/openfactcheck) and publicly released as a Python library (https://pypi.org/project/openfactcheck/) and also as a web service (http://app.openfactcheck.com). A video describing the system is available at https://youtu.be/-i9VKL0HleI.
title OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2408.11832