FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Haoxuan, Liu, Yang, Chong, Bin, Yang, Yingguang, Chen, Yancheng, Liang, Jiayu, Li, Qian, Lu, Hanning, Xu, Kefu, Zheng, Hao, Zhang, Chongyang, Peng, Hao, Yu, Philip S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914605392986112
author Jia, Haoxuan
Liu, Yang
Chong, Bin
Yang, Yingguang
Chen, Yancheng
Liang, Jiayu
Li, Qian
Lu, Hanning
Xu, Kefu
Zheng, Hao
Zhang, Chongyang
Peng, Hao
Yu, Philip S.
author_facet Jia, Haoxuan
Liu, Yang
Chong, Bin
Yang, Yingguang
Chen, Yancheng
Liang, Jiayu
Li, Qian
Lu, Hanning
Xu, Kefu
Zheng, Hao
Zhang, Chongyang
Peng, Hao
Yu, Philip S.
contents Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary filters often miss irreversible mid-trajectory tool calls, while post-hoc LLM judges perform auditing only after termination -- too late for intervention and at a computational cost that scales linearly with trace length. We present FinHarness, an inline safety harness that wraps a finance agent end-to-end with three components: a Query Monitor that fuses single-turn intent with cross-turn drift, a Tool Monitor that evaluates each prospective tool call, and a Cascade module that integrates per-step risk and adaptively routes verification between a lightweight and an advanced-tier LLM judge. Fired risk factors are re-injected into the agent input as ex-ante evidence, enabling the agent to refuse, re-plan, or approve on its own. On FinVault, routed FinHarness cuts ASR from 38.3% to 15.0% while largely preserving benign approval ($41.1\% \to 39.3\%$), and uses $4.7\times$ fewer advanced-judge calls than an always-advanced ablation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27333
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
Jia, Haoxuan
Liu, Yang
Chong, Bin
Yang, Yingguang
Chen, Yancheng
Liang, Jiayu
Li, Qian
Lu, Hanning
Xu, Kefu
Zheng, Hao
Zhang, Chongyang
Peng, Hao
Yu, Philip S.
Computation and Language
Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary filters often miss irreversible mid-trajectory tool calls, while post-hoc LLM judges perform auditing only after termination -- too late for intervention and at a computational cost that scales linearly with trace length. We present FinHarness, an inline safety harness that wraps a finance agent end-to-end with three components: a Query Monitor that fuses single-turn intent with cross-turn drift, a Tool Monitor that evaluates each prospective tool call, and a Cascade module that integrates per-step risk and adaptively routes verification between a lightweight and an advanced-tier LLM judge. Fired risk factors are re-injected into the agent input as ex-ante evidence, enabling the agent to refuse, re-plan, or approve on its own. On FinVault, routed FinHarness cuts ASR from 38.3% to 15.0% while largely preserving benign approval ($41.1\% \to 39.3\%$), and uses $4.7\times$ fewer advanced-judge calls than an always-advanced ablation.
title FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
topic Computation and Language
url https://arxiv.org/abs/2605.27333