FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Mengao, Fu, Jiayu, Warrier, Tanya, Wang, Yuwen, Tan, Tianhui, Huang, Ke-wei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911228934225920
author Zhang, Mengao
Fu, Jiayu
Warrier, Tanya
Wang, Yuwen
Tan, Tianhui
Huang, Ke-wei
author_facet Zhang, Mengao
Fu, Jiayu
Warrier, Tanya
Wang, Yuwen
Tan, Tianhui
Huang, Ke-wei
contents Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for reliable financial analysis, since even minor numerical errors can undermine decision-making and regulatory compliance. Financial applications have unique requirements, often relying on context-dependent, numerical, and proprietary tabular data that existing hallucination benchmarks rarely capture. In this study, we develop a rigorous and scalable framework for evaluating intrinsic hallucinations in financial LLMs, conceptualized as a context-aware masked span prediction task over real-world financial documents. Our main contributions are: (1) a novel, automated dataset creation paradigm using a masking strategy; (2) a new hallucination evaluation dataset derived from S&P 500 annual reports; and (3) a comprehensive evaluation of intrinsic hallucination patterns in state-of-the-art LLMs on financial tabular data. Our work provides a robust methodology for in-house LLM evaluation and serves as a critical step toward building more trustworthy and reliable financial Generative AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05201
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
Zhang, Mengao
Fu, Jiayu
Warrier, Tanya
Wang, Yuwen
Tan, Tianhui
Huang, Ke-wei
Machine Learning
Artificial Intelligence
Computation and Language
Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for reliable financial analysis, since even minor numerical errors can undermine decision-making and regulatory compliance. Financial applications have unique requirements, often relying on context-dependent, numerical, and proprietary tabular data that existing hallucination benchmarks rarely capture. In this study, we develop a rigorous and scalable framework for evaluating intrinsic hallucinations in financial LLMs, conceptualized as a context-aware masked span prediction task over real-world financial documents. Our main contributions are: (1) a novel, automated dataset creation paradigm using a masking strategy; (2) a new hallucination evaluation dataset derived from S&P 500 annual reports; and (3) a comprehensive evaluation of intrinsic hallucination patterns in state-of-the-art LLMs on financial tabular data. Our work provides a robust methodology for in-house LLM evaluation and serves as a critical step toward building more trustworthy and reliable financial Generative AI systems.
title FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.05201