FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wan, Yingjia, Tan, Haochen, Zhu, Xiao, Zhou, Xinyu, Li, Zhiwei, Lv, Qingsong, Sun, Changxuan, Zeng, Jiaqi, Xu, Yi, Lu, Jianqiao, Liu, Yinhong, Guo, Zhijiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917060666195968
author Wan, Yingjia
Tan, Haochen
Zhu, Xiao
Zhou, Xinyu
Li, Zhiwei
Lv, Qingsong
Sun, Changxuan
Zeng, Jiaqi
Xu, Yi
Lu, Jianqiao
Liu, Yinhong
Guo, Zhijiang
author_facet Wan, Yingjia
Tan, Haochen
Zhu, Xiao
Zhou, Xinyu
Li, Zhiwei
Lv, Qingsong
Sun, Changxuan
Zeng, Jiaqi
Xu, Yi
Lu, Jianqiao
Liu, Yinhong
Guo, Zhijiang
contents Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to efficiency bottlenecks and reliability concerns. Prior efforts attempt this by decomposing text into claims, searching for evidence, and verifying claims, but suffer from critical drawbacks: (1) inefficiency due to overcomplicated pipeline components, and (2) ineffectiveness stemming from inaccurate claim sets and insufficient evidence. To address these limitations, we propose \textbf{FaStfact}, an evaluation framework that achieves the highest alignment with human evaluation and time/token efficiency among existing baselines. FaStfact first employs chunk-level claim extraction integrated with confidence-based pre-verification, significantly reducing the time and token cost while ensuring reliability. For searching and verification, it collects document-level evidence from crawled web-pages and selectively retrieves it during verification. Extensive experiments based on an annotated benchmark \textbf{FaStfact-Bench} demonstrate the reliability of FaStfact in both efficiently and effectively evaluating long-form factuality. Code, benchmark data, and annotation interface tool are available at https://github.com/Yingjia-Wan/FaStfact.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
Wan, Yingjia
Tan, Haochen
Zhu, Xiao
Zhou, Xinyu
Li, Zhiwei
Lv, Qingsong
Sun, Changxuan
Zeng, Jiaqi
Xu, Yi
Lu, Jianqiao
Liu, Yinhong
Guo, Zhijiang
Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to efficiency bottlenecks and reliability concerns. Prior efforts attempt this by decomposing text into claims, searching for evidence, and verifying claims, but suffer from critical drawbacks: (1) inefficiency due to overcomplicated pipeline components, and (2) ineffectiveness stemming from inaccurate claim sets and insufficient evidence. To address these limitations, we propose \textbf{FaStfact}, an evaluation framework that achieves the highest alignment with human evaluation and time/token efficiency among existing baselines. FaStfact first employs chunk-level claim extraction integrated with confidence-based pre-verification, significantly reducing the time and token cost while ensuring reliability. For searching and verification, it collects document-level evidence from crawled web-pages and selectively retrieves it during verification. Extensive experiments based on an annotated benchmark \textbf{FaStfact-Bench} demonstrate the reliability of FaStfact in both efficiently and effectively evaluating long-form factuality. Code, benchmark data, and annotation interface tool are available at https://github.com/Yingjia-Wan/FaStfact.
title FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
topic Computation and Language
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
url https://arxiv.org/abs/2510.12839