Beyond Knowledge to Agency: Evaluating Expertise, Autonomy, and Integrity in Finance with CNFinBench

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jinru, Ding, Chao, Jiang, Yidong, Pang, Wenrao, Xiao, Boyi, Liu, Zhiqiang, Chen, Jiayuan, Zhong, Yun, Yuan, Tiantian, Guan, Junming, Cheng, Dawei, Xu, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910271060049920
author Ding, Jinru
Ding, Chao
Jiang, Yidong
Pang, Wenrao
Xiao, Boyi
Liu, Zhiqiang
Chen, Jiayuan
Zhong, Yun
Yuan, Tiantian
Guan, Junming
Cheng, Dawei
Xu, Jie
author_facet Ding, Jinru
Ding, Chao
Jiang, Yidong
Pang, Wenrao
Xiao, Boyi
Liu, Zhiqiang
Chen, Jiayuan
Zhong, Yun
Yuan, Tiantian
Guan, Junming
Cheng, Dawei
Xu, Jie
contents As large language models (LLMs) become high-privilege agents in risk-sensitive settings, they introduce systemic threats beyond hallucination, where minor compliance errors can cause critical data leaks. However, existing benchmarks focus on rule-based QA, lacking agentic execution modeling, overlooking compliance drift in adversarial interactions, and relying on binary safety metrics that fail to capture behavioral degradation. To bridge these gaps, we present CNFinBench, a comprehensive benchmark spanning 29 subtasks grounded in the triad of expertise, autonomy, and integrity. It assesses domain-specific capabilities through certified regulatory corpora and professional financial tasks, reconstructs end-to-end agent workflows from requirement parsing to tool verification, and simulates multi-turn adversarial attacks that induce behavioral compliance drift. To quantify safety degradation, we introduce the Harmful Instruction Compliance Score (HICS), a multi-dimensional safety metric that integrates risk-type-specific deductions, multi-turn consistency tracking, and severity-adjusted penalty scaling based on fine-grained violation triggers. Evaluations over 22 open-/closed-source models reveal: LLMs perform well in applied tasks yet lack robust rule understanding, suffer a 15.4 decline from single modules to full execution chains, and collapse rapidly in multi-turn attacks, with average violations surging by 159.05% in Round 2. CNFinBench is available at https://cnfinbench.opencompass.org.cn and https://github.com/VertiAIBench/CNFinBench.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Knowledge to Agency: Evaluating Expertise, Autonomy, and Integrity in Finance with CNFinBench
Ding, Jinru
Ding, Chao
Jiang, Yidong
Pang, Wenrao
Xiao, Boyi
Liu, Zhiqiang
Chen, Jiayuan
Zhong, Yun
Yuan, Tiantian
Guan, Junming
Cheng, Dawei
Xu, Jie
Computational Engineering, Finance, and Science
As large language models (LLMs) become high-privilege agents in risk-sensitive settings, they introduce systemic threats beyond hallucination, where minor compliance errors can cause critical data leaks. However, existing benchmarks focus on rule-based QA, lacking agentic execution modeling, overlooking compliance drift in adversarial interactions, and relying on binary safety metrics that fail to capture behavioral degradation. To bridge these gaps, we present CNFinBench, a comprehensive benchmark spanning 29 subtasks grounded in the triad of expertise, autonomy, and integrity. It assesses domain-specific capabilities through certified regulatory corpora and professional financial tasks, reconstructs end-to-end agent workflows from requirement parsing to tool verification, and simulates multi-turn adversarial attacks that induce behavioral compliance drift. To quantify safety degradation, we introduce the Harmful Instruction Compliance Score (HICS), a multi-dimensional safety metric that integrates risk-type-specific deductions, multi-turn consistency tracking, and severity-adjusted penalty scaling based on fine-grained violation triggers. Evaluations over 22 open-/closed-source models reveal: LLMs perform well in applied tasks yet lack robust rule understanding, suffer a 15.4 decline from single modules to full execution chains, and collapse rapidly in multi-turn attacks, with average violations surging by 159.05% in Round 2. CNFinBench is available at https://cnfinbench.opencompass.org.cn and https://github.com/VertiAIBench/CNFinBench.
title Beyond Knowledge to Agency: Evaluating Expertise, Autonomy, and Integrity in Finance with CNFinBench
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2512.09506