FinSheet-Bench: From Simple Lookups to Complex Reasoning, Where LLMs Break on Financial Spreadsheets
Fuente:
arXiv
Saved in:
| Main Authors: | Ravnik, Jan, Ličen, Matjaž, Bührmann, Felix, Yuan, Bithiah, Stinson, Felix, Singh, Tanvi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FinBERT-QA: Financial Question Answering with pre-trained BERT Language Models
by: Yuan, Bithiah
Published: (2025)
by: Yuan, Bithiah
Published: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
by: Agrawal, Yogesh, et al.
Published: (2026)
by: Agrawal, Yogesh, et al.
Published: (2026)
SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Data for All-optical programmable liquid-crystal photonic circuits
by: Korenjak, Zala, et al.
Published: (2026)
by: Korenjak, Zala, et al.
Published: (2026)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
by: Ma, Zeyao, et al.
Published: (2024)
by: Ma, Zeyao, et al.
Published: (2024)
FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models
by: Yuan, Ziqiang, et al.
Published: (2024)
by: Yuan, Ziqiang, et al.
Published: (2024)
SheetAgent: Towards A Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models
by: Chen, Yibin, et al.
Published: (2024)
by: Chen, Yibin, et al.
Published: (2024)
Sheet as Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
by: Lei, Yiming, et al.
Published: (2026)
by: Lei, Yiming, et al.
Published: (2026)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
by: Hu, Xuesi, et al.
Published: (2026)
by: Hu, Xuesi, et al.
Published: (2026)
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
by: Malarkkan, Arun Vignesh, et al.
Published: (2026)
by: Malarkkan, Arun Vignesh, et al.
Published: (2026)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
Component parts of the World Heat Flow Data Collection
by: Ravnik, D
Published: (1991)
by: Ravnik, D
Published: (1991)
Disaster Nursing Competencies in a Time of Global Conflicts and Climate Crises: A Cross‐Sectional Survey Study
by: Sabina Ličen, et al.
Published: (2025)
by: Sabina Ličen, et al.
Published: (2025)
Spirituality, Culture and Job Satisfaction in the Holistic Assessment of Nurses’ Well‐Being at Work: A Cross‐Sectional Survey Study
by: Sabina Ličen, et al.
Published: (2025)
by: Sabina Ličen, et al.
Published: (2025)
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
by: Pall, Rajveer Singh
Published: (2026)
by: Pall, Rajveer Singh
Published: (2026)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
by: Lu, Guilong, et al.
Published: (2025)
by: Lu, Guilong, et al.
Published: (2025)
Reasoning Fails Where Step Flow Breaks
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Controls over Spreadsheets for Financial Reporting in Practice
by: Coster, Nancy, et al.
Published: (2011)
by: Coster, Nancy, et al.
Published: (2011)
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains
by: Park, Seunghyun, et al.
Published: (2026)
by: Park, Seunghyun, et al.
Published: (2026)
FinCARE: Financial Causal Analysis with Reasoning and Evidence
by: Michel, Alejandro, et al.
Published: (2025)
by: Michel, Alejandro, et al.
Published: (2025)
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
by: Zhu, Ruiyan, et al.
Published: (2025)
by: Zhu, Ruiyan, et al.
Published: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
by: Hou, Yutao, et al.
Published: (2026)
by: Hou, Yutao, et al.
Published: (2026)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection
by: Chen, Qin, et al.
Published: (2025)
by: Chen, Qin, et al.
Published: (2025)
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
by: Nitarach, Natapong, et al.
Published: (2025)
by: Nitarach, Natapong, et al.
Published: (2025)
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
by: Dimino, Fabrizio, et al.
Published: (2025)
by: Dimino, Fabrizio, et al.
Published: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
by: Lu, Jiaxuan, et al.
Published: (2026)
by: Lu, Jiaxuan, et al.
Published: (2026)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
by: Choi, Chanyeol, et al.
Published: (2025)
by: Choi, Chanyeol, et al.
Published: (2025)
From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding
by: Gulati, Anmol, et al.
Published: (2026)
by: Gulati, Anmol, et al.
Published: (2026)
Demystifying Diffusion Policies: Action Memorization and Simple Lookup Table Alternatives
by: He, Chengyang, et al.
Published: (2025)
by: He, Chengyang, et al.
Published: (2025)
FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models
by: Lee, Seunghan, et al.
Published: (2026)
by: Lee, Seunghan, et al.
Published: (2026)
Direct Air Capture in Europe - Where to Integrate, Where to Store, and What Drives Cost?
by: Bernecker, Maximilian, et al.
Published: (2026)
by: Bernecker, Maximilian, et al.
Published: (2026)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025)
by: Xie, Zhuohan, et al.
Published: (2025)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
by: El-Haj, Mo, et al.
Published: (2025)
by: El-Haj, Mo, et al.
Published: (2025)
Open FinLLM Leaderboard: Towards Financial AI Readiness
by: Lin, Shengyuan Colin, et al.
Published: (2025)
by: Lin, Shengyuan Colin, et al.
Published: (2025)
FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Statement Verification
by: Panda, Silu
Published: (2026)
by: Panda, Silu
Published: (2026)
Similar Items
-
FinBERT-QA: Financial Question Answering with pre-trained BERT Language Models
by: Yuan, Bithiah
Published: (2025) -
FinTradeBench: A Financial Reasoning Benchmark for LLMs
by: Agrawal, Yogesh, et al.
Published: (2026) -
SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets
by: Wang, Ziwei, et al.
Published: (2025) -
Data for All-optical programmable liquid-crystal photonic circuits
by: Korenjak, Zala, et al.
Published: (2026) -
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)