All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912810588438528 |
|---|---|
| author | Jiang, Yuechen Liu, Zhiwei Cao, Yupeng He, Yueru Xu, Ziyang Xu, Chen Deng, Zhiyang Tiwari, Prayag Chen, Xi Lopez-Lira, Alejandro Huang, Jimin Tsujii, Junichi Ananiadou, Sophia |
| author_facet | Jiang, Yuechen Liu, Zhiwei Cao, Yupeng He, Yueru Xu, Ziyang Xu, Chen Deng, Zhiyang Tiwari, Prayag Chen, Xi Lopez-Lira, Alejandro Huang, Jimin Tsujii, Junichi Ananiadou, Sophia |
| contents | We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures the contextual complexity of financial news where meaning emerges from dispersed cues. The benchmark defines two complementary tasks: reference free misinformation detection and comparison based diagnosis using paired original perturbed inputs. Experiments reveal a consistent pattern: performance is substantially stronger when comparative context is available, while reference free settings expose significant weaknesses, including unstable predictions and elevated invalid outputs. These results indicate that current models struggle to maintain coherent belief states without external grounding. By highlighting this gap, RFC Bench provides a structured testbed for studying reference free reasoning and advancing more reliable financial misinformation detection in real world settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_04160 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection Jiang, Yuechen Liu, Zhiwei Cao, Yupeng He, Yueru Xu, Ziyang Xu, Chen Deng, Zhiyang Tiwari, Prayag Chen, Xi Lopez-Lira, Alejandro Huang, Jimin Tsujii, Junichi Ananiadou, Sophia Computation and Language Computational Engineering, Finance, and Science Computational Finance We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures the contextual complexity of financial news where meaning emerges from dispersed cues. The benchmark defines two complementary tasks: reference free misinformation detection and comparison based diagnosis using paired original perturbed inputs. Experiments reveal a consistent pattern: performance is substantially stronger when comparative context is available, while reference free settings expose significant weaknesses, including unstable predictions and elevated invalid outputs. These results indicate that current models struggle to maintain coherent belief states without external grounding. By highlighting this gap, RFC Bench provides a structured testbed for studying reference free reasoning and advancing more reliable financial misinformation detection in real world settings. |
| title | All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection |
| topic | Computation and Language Computational Engineering, Finance, and Science Computational Finance |
| url | https://arxiv.org/abs/2601.04160 |