FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917060666195968 |
|---|---|
| author | Wan, Yingjia Tan, Haochen Zhu, Xiao Zhou, Xinyu Li, Zhiwei Lv, Qingsong Sun, Changxuan Zeng, Jiaqi Xu, Yi Lu, Jianqiao Liu, Yinhong Guo, Zhijiang |
| author_facet | Wan, Yingjia Tan, Haochen Zhu, Xiao Zhou, Xinyu Li, Zhiwei Lv, Qingsong Sun, Changxuan Zeng, Jiaqi Xu, Yi Lu, Jianqiao Liu, Yinhong Guo, Zhijiang |
| contents | Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to efficiency bottlenecks and reliability concerns. Prior efforts attempt this by decomposing text into claims, searching for evidence, and verifying claims, but suffer from critical drawbacks: (1) inefficiency due to overcomplicated pipeline components, and (2) ineffectiveness stemming from inaccurate claim sets and insufficient evidence. To address these limitations, we propose \textbf{FaStfact}, an evaluation framework that achieves the highest alignment with human evaluation and time/token efficiency among existing baselines. FaStfact first employs chunk-level claim extraction integrated with confidence-based pre-verification, significantly reducing the time and token cost while ensuring reliability. For searching and verification, it collects document-level evidence from crawled web-pages and selectively retrieves it during verification. Extensive experiments based on an annotated benchmark \textbf{FaStfact-Bench} demonstrate the reliability of FaStfact in both efficiently and effectively evaluating long-form factuality. Code, benchmark data, and annotation interface tool are available at https://github.com/Yingjia-Wan/FaStfact. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_12839 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs Wan, Yingjia Tan, Haochen Zhu, Xiao Zhou, Xinyu Li, Zhiwei Lv, Qingsong Sun, Changxuan Zeng, Jiaqi Xu, Yi Lu, Jianqiao Liu, Yinhong Guo, Zhijiang Computation and Language Artificial Intelligence Computational Engineering, Finance, and Science Computers and Society Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to efficiency bottlenecks and reliability concerns. Prior efforts attempt this by decomposing text into claims, searching for evidence, and verifying claims, but suffer from critical drawbacks: (1) inefficiency due to overcomplicated pipeline components, and (2) ineffectiveness stemming from inaccurate claim sets and insufficient evidence. To address these limitations, we propose \textbf{FaStfact}, an evaluation framework that achieves the highest alignment with human evaluation and time/token efficiency among existing baselines. FaStfact first employs chunk-level claim extraction integrated with confidence-based pre-verification, significantly reducing the time and token cost while ensuring reliability. For searching and verification, it collects document-level evidence from crawled web-pages and selectively retrieves it during verification. Extensive experiments based on an annotated benchmark \textbf{FaStfact-Bench} demonstrate the reliability of FaStfact in both efficiently and effectively evaluating long-form factuality. Code, benchmark data, and annotation interface tool are available at https://github.com/Yingjia-Wan/FaStfact. |
| title | FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs |
| topic | Computation and Language Artificial Intelligence Computational Engineering, Finance, and Science Computers and Society |
| url | https://arxiv.org/abs/2510.12839 |