RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dai, Yuyang, Lin, Yan, Xie, Zhuohan, Wang, Yuxia
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917435984052224
author Dai, Yuyang
Lin, Yan
Xie, Zhuohan
Wang, Yuxia
author_facet Dai, Yuyang
Lin, Yan
Xie, Zhuohan
Wang, Yuxia
contents Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice, problems often rely on implicit assumptions that are taken for granted rather than stated explicitly, causing problems to appear solvable while lacking enough information for a definite answer. We introduce REALFIN, a bilingual benchmark that evaluates financial reasoning by systematically removing essential premises from exam-style questions while keeping them linguistically plausible. Based on this, we evaluate models under three formulations that test answering, recognizing missing information, and rejecting unjustified options, and find consistent performance drops when key conditions are absent. General-purpose models tend to over-commit and guess, while most finance-specialized models fail to clearly identify missing premises. These results highlight a critical gap in current evaluations and show that reliable financial models must know when a question should not be answered.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07096
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
Dai, Yuyang
Lin, Yan
Xie, Zhuohan
Wang, Yuxia
Statistical Finance
Artificial Intelligence
Computational Finance
Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice, problems often rely on implicit assumptions that are taken for granted rather than stated explicitly, causing problems to appear solvable while lacking enough information for a definite answer. We introduce REALFIN, a bilingual benchmark that evaluates financial reasoning by systematically removing essential premises from exam-style questions while keeping them linguistically plausible. Based on this, we evaluate models under three formulations that test answering, recognizing missing information, and rejecting unjustified options, and find consistent performance drops when key conditions are absent. General-purpose models tend to over-commit and guess, while most finance-specialized models fail to clearly identify missing premises. These results highlight a critical gap in current evaluations and show that reliable financial models must know when a question should not be answered.
title RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
topic Statistical Finance
Artificial Intelligence
Computational Finance
url https://arxiv.org/abs/2602.07096