VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tai, Zhenghan, Wu, Hanwei, Hu, Qingchen, Chi, Jijun, He, Hailin, Ding, Lei, Kwok, Tung Sum Thomas, Xiao, Bohuai, Hua, Yuchen, Wang, Suyuchen, Lu, Peng, Li, Muzhi, Wu, Yihong, Ma, Liheng, Huang, Jerry, Zhang, Jiayi, Zhang, Gonghao, Jiang, Chaolong, Tian, Jingrui, Lyu, Sicheng, Li, Zeyu, Han, Boyu, Mo, Fengran, Yu, Xinyue, Cui, Yufei, Zhou, Ling, Wang, Xinyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912644222418944
author Tai, Zhenghan
Wu, Hanwei
Hu, Qingchen
Chi, Jijun
He, Hailin
Ding, Lei
Kwok, Tung Sum Thomas
Xiao, Bohuai
Hua, Yuchen
Wang, Suyuchen
Lu, Peng
Li, Muzhi
Wu, Yihong
Ma, Liheng
Huang, Jerry
Zhang, Jiayi
Zhang, Gonghao
Jiang, Chaolong
Tian, Jingrui
Lyu, Sicheng
Li, Zeyu
Han, Boyu
Mo, Fengran
Yu, Xinyue
Cui, Yufei
Zhou, Ling
Wang, Xinyu
author_facet Tai, Zhenghan
Wu, Hanwei
Hu, Qingchen
Chi, Jijun
He, Hailin
Ding, Lei
Kwok, Tung Sum Thomas
Xiao, Bohuai
Hua, Yuchen
Wang, Suyuchen
Lu, Peng
Li, Muzhi
Wu, Yihong
Ma, Liheng
Huang, Jerry
Zhang, Jiayi
Zhang, Gonghao
Jiang, Chaolong
Tian, Jingrui
Lyu, Sicheng
Li, Zeyu
Han, Boyu
Mo, Fengran
Yu, Xinyue
Cui, Yufei
Zhou, Ling
Wang, Xinyu
contents Retrieval-Augmented Generation (RAG) is becoming increasingly essential for Question Answering (QA) in the financial sector, where accurate and contextually grounded insights from complex public disclosures are crucial. However, existing financial RAG systems face two significant challenges: (1) they struggle to process heterogeneous data formats, such as text, tables, and figures; and (2) they encounter difficulties in balancing general-domain applicability with company-specific adaptation. To overcome these challenges, we present VeritasFi, an innovative hybrid RAG framework that incorporates a multi-modal preprocessing pipeline alongside a cutting-edge two-stage training strategy for its re-ranking component. VeritasFi enhances financial QA through three key innovations: (1) A multi-modal preprocessing pipeline that seamlessly transforms heterogeneous data into a coherent, machine-readable format. (2) A tripartite hybrid retrieval engine that operates in parallel, combining deep multi-path retrieval over a semantically indexed document corpus, real-time data acquisition through tool utilization, and an expert-curated memory bank for high-frequency questions, ensuring comprehensive scope, accuracy, and efficiency. (3) A two-stage training strategy for the document re-ranker, which initially constructs a general, domain-specific model using anonymized data, followed by rapid fine-tuning on company-specific data for targeted applications. By integrating our proposed designs, VeritasFi presents a groundbreaking framework that greatly enhances the adaptability and robustness of financial RAG systems, providing a scalable solution for both general-domain and company-specific QA tasks. Code accompanying this work is available at https://github.com/simplew4y/VeritasFi.git.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
Tai, Zhenghan
Wu, Hanwei
Hu, Qingchen
Chi, Jijun
He, Hailin
Ding, Lei
Kwok, Tung Sum Thomas
Xiao, Bohuai
Hua, Yuchen
Wang, Suyuchen
Lu, Peng
Li, Muzhi
Wu, Yihong
Ma, Liheng
Huang, Jerry
Zhang, Jiayi
Zhang, Gonghao
Jiang, Chaolong
Tian, Jingrui
Lyu, Sicheng
Li, Zeyu
Han, Boyu
Mo, Fengran
Yu, Xinyue
Cui, Yufei
Zhou, Ling
Wang, Xinyu
Information Retrieval
Artificial Intelligence
Retrieval-Augmented Generation (RAG) is becoming increasingly essential for Question Answering (QA) in the financial sector, where accurate and contextually grounded insights from complex public disclosures are crucial. However, existing financial RAG systems face two significant challenges: (1) they struggle to process heterogeneous data formats, such as text, tables, and figures; and (2) they encounter difficulties in balancing general-domain applicability with company-specific adaptation. To overcome these challenges, we present VeritasFi, an innovative hybrid RAG framework that incorporates a multi-modal preprocessing pipeline alongside a cutting-edge two-stage training strategy for its re-ranking component. VeritasFi enhances financial QA through three key innovations: (1) A multi-modal preprocessing pipeline that seamlessly transforms heterogeneous data into a coherent, machine-readable format. (2) A tripartite hybrid retrieval engine that operates in parallel, combining deep multi-path retrieval over a semantically indexed document corpus, real-time data acquisition through tool utilization, and an expert-curated memory bank for high-frequency questions, ensuring comprehensive scope, accuracy, and efficiency. (3) A two-stage training strategy for the document re-ranker, which initially constructs a general, domain-specific model using anonymized data, followed by rapid fine-tuning on company-specific data for targeted applications. By integrating our proposed designs, VeritasFi presents a groundbreaking framework that greatly enhances the adaptability and robustness of financial RAG systems, providing a scalable solution for both general-domain and company-specific QA tasks. Code accompanying this work is available at https://github.com/simplew4y/VeritasFi.git.
title VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2510.10828