QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shupeng, Lu, Weipeng, Liu, Linyun, Lin, Chen, Li, Shaofei, Tan, Zhendong, Zhong, Hanjun, Zeng, Yucheng, Zhu, Chenghao, Liu, Mengyue, Dong, Daxiang, Wu, Jianmin, Xiao, Yunting, Li, Annan, Liu, Danyu, Zhang, Jingnan, Liu, Licen, Yin, Dawei, Shen, Dou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918294564372480
author Li, Shupeng
Lu, Weipeng
Liu, Linyun
Lin, Chen
Li, Shaofei
Tan, Zhendong
Zhong, Hanjun
Zeng, Yucheng
Zhu, Chenghao
Liu, Mengyue
Dong, Daxiang
Wu, Jianmin
Xiao, Yunting
Li, Annan
Liu, Danyu
Zhang, Jingnan
Liu, Licen
Yin, Dawei
Shen, Dou
author_facet Li, Shupeng
Lu, Weipeng
Liu, Linyun
Lin, Chen
Li, Shaofei
Tan, Zhendong
Zhong, Hanjun
Zeng, Yucheng
Zhu, Chenghao
Liu, Mengyue
Dong, Daxiang
Wu, Jianmin
Xiao, Yunting
Li, Annan
Liu, Danyu
Zhang, Jingnan
Liu, Licen
Yin, Dawei
Shen, Dou
contents Domain-specific enhancement of Large Language Models (LLMs) within the financial context has long been a focal point of industrial application. While previous models such as BloombergGPT and Baichuan-Finance primarily focused on knowledge enhancement, the deepening complexity of financial services has driven a growing demand for models that possess not only domain knowledge but also robust financial reasoning and agentic capabilities. In this paper, we present QianfanHuijin, a financial domain LLM, and propose a generalizable multi-stage training paradigm for industrial model enhancement. Our approach begins with Continual Pre-training (CPT) on financial corpora to consolidate the knowledge base. This is followed by a fine-grained Post-training pipeline designed with increasing specificity: starting with Financial SFT, progressing to Finance Reasoning RL and Finance Agentic RL, and culminating in General RL aligned with real-world business scenarios. Empirical results demonstrate that QianfanHuijin achieves superior performance across various authoritative financial benchmarks. Furthermore, ablation studies confirm that the targeted Reasoning RL and Agentic RL stages yield significant gains in their respective capabilities. These findings validate our motivation and suggest that this fine-grained, progressive post-training methodology is poised to become a mainstream paradigm for various industrial-enhanced LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
Li, Shupeng
Lu, Weipeng
Liu, Linyun
Lin, Chen
Li, Shaofei
Tan, Zhendong
Zhong, Hanjun
Zeng, Yucheng
Zhu, Chenghao
Liu, Mengyue
Dong, Daxiang
Wu, Jianmin
Xiao, Yunting
Li, Annan
Liu, Danyu
Zhang, Jingnan
Liu, Licen
Yin, Dawei
Shen, Dou
Computation and Language
Domain-specific enhancement of Large Language Models (LLMs) within the financial context has long been a focal point of industrial application. While previous models such as BloombergGPT and Baichuan-Finance primarily focused on knowledge enhancement, the deepening complexity of financial services has driven a growing demand for models that possess not only domain knowledge but also robust financial reasoning and agentic capabilities. In this paper, we present QianfanHuijin, a financial domain LLM, and propose a generalizable multi-stage training paradigm for industrial model enhancement. Our approach begins with Continual Pre-training (CPT) on financial corpora to consolidate the knowledge base. This is followed by a fine-grained Post-training pipeline designed with increasing specificity: starting with Financial SFT, progressing to Finance Reasoning RL and Finance Agentic RL, and culminating in General RL aligned with real-world business scenarios. Empirical results demonstrate that QianfanHuijin achieves superior performance across various authoritative financial benchmarks. Furthermore, ablation studies confirm that the targeted Reasoning RL and Agentic RL stages yield significant gains in their respective capabilities. These findings validate our motivation and suggest that this fine-grained, progressive post-training methodology is poised to become a mainstream paradigm for various industrial-enhanced LLMs.
title QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
topic Computation and Language
url https://arxiv.org/abs/2512.24314