Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choe, Jaeyoung, Kim, Jihoon, Jung, Woohwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915601100832768
author Choe, Jaeyoung
Kim, Jihoon
Jung, Woohwan
author_facet Choe, Jaeyoung
Kim, Jihoon
Jung, Woohwan
contents Retrieval-augmented generation (RAG) based large language models (LLMs) are widely used in finance for their excellent performance on knowledge-intensive tasks. However, standardized documents (e.g., SEC filing) share similar formats such as repetitive boilerplate texts, and similar table structures. This similarity forces traditional RAG methods to misidentify near-duplicate text, leading to duplicate retrieval that undermines accuracy and completeness. To address these issues, we propose the Hierarchical Retrieval with Evidence Curation (HiREC) framework. Our approach first performs hierarchical retrieval to reduce confusion among similar texts. It first retrieve related documents and then selects the most relevant passages from the documents. The evidence curation process removes irrelevant passages. When necessary, it automatically generates complementary queries to collect missing information. To evaluate our approach, we construct and release a Large-scale Open-domain Financial (LOFin) question answering benchmark that includes 145,897 SEC documents and 1,595 question-answer pairs. Our code and data are available at https://github.com/deep-over/LOFin-bench-HiREC.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
Choe, Jaeyoung
Kim, Jihoon
Jung, Woohwan
Information Retrieval
Artificial Intelligence
Computation and Language
Retrieval-augmented generation (RAG) based large language models (LLMs) are widely used in finance for their excellent performance on knowledge-intensive tasks. However, standardized documents (e.g., SEC filing) share similar formats such as repetitive boilerplate texts, and similar table structures. This similarity forces traditional RAG methods to misidentify near-duplicate text, leading to duplicate retrieval that undermines accuracy and completeness. To address these issues, we propose the Hierarchical Retrieval with Evidence Curation (HiREC) framework. Our approach first performs hierarchical retrieval to reduce confusion among similar texts. It first retrieve related documents and then selects the most relevant passages from the documents. The evidence curation process removes irrelevant passages. When necessary, it automatically generates complementary queries to collect missing information. To evaluate our approach, we construct and release a Large-scale Open-domain Financial (LOFin) question answering benchmark that includes 145,897 SEC documents and 1,595 question-answer pairs. Our code and data are available at https://github.com/deep-over/LOFin-bench-HiREC.
title Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.20368