Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yukun, Droguett, Stefan Elbl, Jain, Samyak
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909978090012672
author Zhang, Yukun
Droguett, Stefan Elbl
Jain, Samyak
author_facet Zhang, Yukun
Droguett, Stefan Elbl
Jain, Samyak
contents This research project addresses the errors of financial numerical reasoning Question Answering (QA) tasks due to the lack of domain knowledge in finance. Despite recent advances in Large Language Models (LLMs), financial numerical questions remain challenging because they require specific domain knowledge in finance and complex multi-step numeric reasoning. We implement a multi-retriever Retrieval Augmented Generators (RAG) system to retrieve both external domain knowledge and internal question contexts, and utilize the latest LLM to tackle these tasks. Through comprehensive ablation experiments and error analysis, we find that domain-specific training with the SecBERT encoder significantly contributes to our best neural symbolic model surpassing the FinQA paper's top model, which serves as our baseline. This suggests the potential superior performance of domain-specific training. Furthermore, our best prompt-based LLM generator achieves the state-of-the-art (SOTA) performance with significant improvement (>7%), yet it is still below the human expert performance. This study highlights the trade-off between hallucinations loss and external knowledge gains in smaller models and few-shot examples. For larger models, the gains from external facts typically outweigh the hallucination loss. Finally, our findings confirm the enhanced numerical reasoning capabilities of the latest LLM, optimized for few-shot learning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs
Zhang, Yukun
Droguett, Stefan Elbl
Jain, Samyak
Computation and Language
Computational Engineering, Finance, and Science
Machine Learning
This research project addresses the errors of financial numerical reasoning Question Answering (QA) tasks due to the lack of domain knowledge in finance. Despite recent advances in Large Language Models (LLMs), financial numerical questions remain challenging because they require specific domain knowledge in finance and complex multi-step numeric reasoning. We implement a multi-retriever Retrieval Augmented Generators (RAG) system to retrieve both external domain knowledge and internal question contexts, and utilize the latest LLM to tackle these tasks. Through comprehensive ablation experiments and error analysis, we find that domain-specific training with the SecBERT encoder significantly contributes to our best neural symbolic model surpassing the FinQA paper's top model, which serves as our baseline. This suggests the potential superior performance of domain-specific training. Furthermore, our best prompt-based LLM generator achieves the state-of-the-art (SOTA) performance with significant improvement (>7%), yet it is still below the human expert performance. This study highlights the trade-off between hallucinations loss and external knowledge gains in smaller models and few-shot examples. For larger models, the gains from external facts typically outweigh the hallucination loss. Finally, our findings confirm the enhanced numerical reasoning capabilities of the latest LLM, optimized for few-shot learning.
title Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs
topic Computation and Language
Computational Engineering, Finance, and Science
Machine Learning
url https://arxiv.org/abs/2512.23848