FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dimino, Fabrizio, Arun, Abhinav, Sarmah, Bhaskarjit, Pasquali, Stefano |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FinReflectKG: Agentic Construction and Evaluation of Financial Knowledge Graphs
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
von: Kumar, Mahesh, et al.
Veröffentlicht: (2026)
von: Kumar, Mahesh, et al.
Veröffentlicht: (2026)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2026)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2026)
FinCARE: Financial Causal Analysis with Reasoning and Evidence
von: Michel, Alejandro, et al.
Veröffentlicht: (2025)
von: Michel, Alejandro, et al.
Veröffentlicht: (2025)
Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2025)
FINCH: Financial Intelligence using Natural language for Contextualized SQL Handling
von: Singh, Avinash Kumar, et al.
Veröffentlicht: (2025)
von: Singh, Avinash Kumar, et al.
Veröffentlicht: (2025)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2026)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
Enhanced Local Explainability and Trust Scores with Random Forest Proximities
von: Rosaler, Joshua, et al.
Veröffentlicht: (2023)
von: Rosaler, Joshua, et al.
Veröffentlicht: (2023)
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
von: Jiang, Yidong, et al.
Veröffentlicht: (2026)
von: Jiang, Yidong, et al.
Veröffentlicht: (2026)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
FinZero: Launching Multi-modal Financial Time Series Forecast with Large Reasoning Model
von: Wang, Yanlong, et al.
Veröffentlicht: (2025)
von: Wang, Yanlong, et al.
Veröffentlicht: (2025)
FinBen: A Holistic Financial Benchmark for Large Language Models
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
von: Wang, Yan, et al.
Veröffentlicht: (2026)
von: Wang, Yan, et al.
Veröffentlicht: (2026)
How to Choose a Threshold for an Evaluation Metric for Large Language Models
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
RNA-KG: An ontology-based knowledge graph for representing interactions involving RNA molecules
von: Cavalleri, Emanuele, et al.
Veröffentlicht: (2023)
von: Cavalleri, Emanuele, et al.
Veröffentlicht: (2023)
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
von: Xie, Zhuohan, et al.
Veröffentlicht: (2026)
von: Xie, Zhuohan, et al.
Veröffentlicht: (2026)
QuantBench: Benchmarking AI Methods for Quantitative Investment
von: Wang, Saizhuo, et al.
Veröffentlicht: (2025)
von: Wang, Saizhuo, et al.
Veröffentlicht: (2025)
FinCast: A Foundation Model for Financial Time-Series Forecasting
von: Zhu, Zhuohang, et al.
Veröffentlicht: (2025)
von: Zhu, Zhuohang, et al.
Veröffentlicht: (2025)
FinSight: Towards Real-World Financial Deep Research
von: Jin, Jiajie, et al.
Veröffentlicht: (2025)
von: Jin, Jiajie, et al.
Veröffentlicht: (2025)
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
von: Li, Haohang, et al.
Veröffentlicht: (2024)
von: Li, Haohang, et al.
Veröffentlicht: (2024)
FinHEAR: Human Expertise and Adaptive Risk-Aware Temporal Reasoning for Financial Decision-Making
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2025)
RiskLabs: Predicting Financial Risk Using Large Language Model based on Multimodal and Multi-Sources Data
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
CatMemo at the FinLLM Challenge Task: Fine-Tuning Large Language Models using Data Fusion in Financial Applications
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
von: Lu, Yuxing, et al.
Veröffentlicht: (2025)
FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
von: Jiang, Junzhe, et al.
Veröffentlicht: (2025)
von: Jiang, Junzhe, et al.
Veröffentlicht: (2025)
SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications
von: Goel, Abhinav, et al.
Veröffentlicht: (2026)
von: Goel, Abhinav, et al.
Veröffentlicht: (2026)
AlphaAgents: Large Language Model based Multi-Agents for Equity Portfolio Constructions
von: Zhao, Tianjiao, et al.
Veröffentlicht: (2025)
von: Zhao, Tianjiao, et al.
Veröffentlicht: (2025)
RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
von: Dai, Yuyang, et al.
Veröffentlicht: (2026)
von: Dai, Yuyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FinReflectKG: Agentic Construction and Evaluation of Financial Knowledge Graphs
von: Arun, Abhinav, et al.
Veröffentlicht: (2025) -
FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
von: Arun, Abhinav, et al.
Veröffentlicht: (2025) -
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
von: Kumar, Mahesh, et al.
Veröffentlicht: (2026) -
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2026) -
FinCARE: Financial Causal Analysis with Reasoning and Evidence
von: Michel, Alejandro, et al.
Veröffentlicht: (2025)