FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Yitao, Hu, Tiansheng, Zhao, Yilun, Cohan, Arman, Zhao, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911197227384832
author Long, Yitao
Hu, Tiansheng
Zhao, Yilun
Cohan, Arman
Zhao, Chen
author_facet Long, Yitao
Hu, Tiansheng
Zhao, Yilun
Cohan, Arman
Zhao, Chen
contents Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribution to LLM outputs. However, existing benchmarks primarily focus on simple attribution that retrieves supporting textual evidence as references. We argue that in real-world scenarios such as financial applications, attribution goes beyond reference retrieval. We introduce FinLFQA, a benchmark designed to evaluate the ability of LLMs to generate long-form answers to complex financial questions with reliable and nuanced attributions. FinLFQA evaluates three critical aspects of attribution through human annotations: (1) supporting evidence extracted from financial reports, (2) intermediate numerical reasoning steps, and (3) domain-specific financial knowledge that informs the reasoning process. We further provide an automatic evaluation framework covering both answer quality and attribution quality. Through extensive experiments on eight LLMs across multiple attribution-generation paradigms, we find that fine-grained metrics are important to distinguish model capabilities, that end-to-end generation achieves comparable performance to post-hoc approaches, and that iterative refinement only helps when guided by external feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
Long, Yitao
Hu, Tiansheng
Zhao, Yilun
Cohan, Arman
Zhao, Chen
Computation and Language
Large Language Models (LLMs) frequently hallucinate to long-form questions, producing plausible yet factually incorrect answers. A common mitigation strategy is to provide attribution to LLM outputs. However, existing benchmarks primarily focus on simple attribution that retrieves supporting textual evidence as references. We argue that in real-world scenarios such as financial applications, attribution goes beyond reference retrieval. We introduce FinLFQA, a benchmark designed to evaluate the ability of LLMs to generate long-form answers to complex financial questions with reliable and nuanced attributions. FinLFQA evaluates three critical aspects of attribution through human annotations: (1) supporting evidence extracted from financial reports, (2) intermediate numerical reasoning steps, and (3) domain-specific financial knowledge that informs the reasoning process. We further provide an automatic evaluation framework covering both answer quality and attribution quality. Through extensive experiments on eight LLMs across multiple attribution-generation paradigms, we find that fine-grained metrics are important to distinguish model capabilities, that end-to-end generation achieves comparable performance to post-hoc approaches, and that iterative refinement only helps when guided by external feedback.
title FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
topic Computation and Language
url https://arxiv.org/abs/2510.06426