FactLens: Benchmarking Fine-Grained Fact Verification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mitra, Kushan, Zhang, Dan, Rahman, Sajjadur, Hruschka, Estevam
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908387145416704
author Mitra, Kushan
Zhang, Dan
Rahman, Sajjadur
Hruschka, Estevam
author_facet Mitra, Kushan
Zhang, Dan
Rahman, Sajjadur
Hruschka, Estevam
contents Large Language Models (LLMs) have shown impressive capability in language generation and understanding, but their tendency to hallucinate and produce factually incorrect information remains a key limitation. To verify LLM-generated contents and claims from other sources, traditional verification approaches often rely on holistic models that assign a single factuality label to complex claims, potentially obscuring nuanced errors. In this paper, we advocate for a shift towards fine-grained verification, where complex claims are broken down into smaller sub-claims for individual verification, allowing for more precise identification of inaccuracies, improved transparency, and reduced ambiguity in evidence retrieval. However, generating sub-claims poses challenges, such as maintaining context and ensuring semantic equivalence with respect to the original claim. We introduce FactLens, a benchmark for evaluating fine-grained fact verification, with metrics and automated evaluators of sub-claim quality. The benchmark data is manually curated to ensure high-quality ground truth. Our results show alignment between automated FactLens evaluators and human judgments, and we discuss the impact of sub-claim characteristics on the overall verification performance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05980
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FactLens: Benchmarking Fine-Grained Fact Verification
Mitra, Kushan
Zhang, Dan
Rahman, Sajjadur
Hruschka, Estevam
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) have shown impressive capability in language generation and understanding, but their tendency to hallucinate and produce factually incorrect information remains a key limitation. To verify LLM-generated contents and claims from other sources, traditional verification approaches often rely on holistic models that assign a single factuality label to complex claims, potentially obscuring nuanced errors. In this paper, we advocate for a shift towards fine-grained verification, where complex claims are broken down into smaller sub-claims for individual verification, allowing for more precise identification of inaccuracies, improved transparency, and reduced ambiguity in evidence retrieval. However, generating sub-claims poses challenges, such as maintaining context and ensuring semantic equivalence with respect to the original claim. We introduce FactLens, a benchmark for evaluating fine-grained fact verification, with metrics and automated evaluators of sub-claim quality. The benchmark data is manually curated to ensure high-quality ground truth. Our results show alignment between automated FactLens evaluators and human judgments, and we discuss the impact of sub-claim characteristics on the overall verification performance.
title FactLens: Benchmarking Fine-Grained Fact Verification
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.05980