FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Xuesi, Wang, Peng, Miao, Jinpeng, Tao, Xilin, Li, Caiwei, Ma, Yue, He, Jie, Zhang, Qiancheng, Zou, Yuntao, Li, Dagang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Factors Affect LLMs and RLLMs in Financial Question Answering?
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
von: Gevers, Ine, et al.
Veröffentlicht: (2025)
von: Gevers, Ine, et al.
Veröffentlicht: (2025)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei
von: Wang, Peng, et al.
Veröffentlicht: (2026)
von: Wang, Peng, et al.
Veröffentlicht: (2026)
Demonstrating SIMA-Play: A Serious Game for Forest Management Decision-Making through Board Game and Digital Simulation
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bi, Zhenyu, et al.
Veröffentlicht: (2025)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
von: Wang, Haochuan, et al.
Veröffentlicht: (2024)
von: Wang, Haochuan, et al.
Veröffentlicht: (2024)
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
von: Kumar, Mahesh, et al.
Veröffentlicht: (2026)
von: Kumar, Mahesh, et al.
Veröffentlicht: (2026)
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayong, et al.
Veröffentlicht: (2026)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhang, Chenxi, et al.
Veröffentlicht: (2026)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
FinBen: A Holistic Financial Benchmark for Large Language Models
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
von: Hou, Yutao, et al.
Veröffentlicht: (2026)
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2025)
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
von: Jiang, Yidong, et al.
Veröffentlicht: (2026)
von: Jiang, Yidong, et al.
Veröffentlicht: (2026)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
Skill vs. Chance Quantification for Popular Card & Board Games
von: Banerjee, Tathagata, et al.
Veröffentlicht: (2024)
von: Banerjee, Tathagata, et al.
Veröffentlicht: (2024)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
Multi-Agent Strategic Games with LLMs
von: Chupilkin, Maxim
Veröffentlicht: (2026)
von: Chupilkin, Maxim
Veröffentlicht: (2026)
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
von: Nitarach, Natapong, et al.
Veröffentlicht: (2025)
von: Nitarach, Natapong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Factors Affect LLMs and RLLMs in Financial Question Answering?
von: Wang, Peng, et al.
Veröffentlicht: (2025) -
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026) -
Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
von: Gevers, Ine, et al.
Veröffentlicht: (2025) -
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025) -
Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei
von: Wang, Peng, et al.
Veröffentlicht: (2026)