Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
Fuente:
arXiv
Guardado en:
| Autores principales: | Haque, Mirazul, Papadimitriou, Antony, Mensah, Samuel, Ma, Zhiqiang, Guo, Zhijin, Sain, Joy Prakash, Kaur, Simerjot, Smiley, Charese, Liu, Xiaomo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
por: Singh, Kuldeep, et al.
Publicado: (2024)
por: Singh, Kuldeep, et al.
Publicado: (2024)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
por: Magomere, Jabez, et al.
Publicado: (2025)
por: Magomere, Jabez, et al.
Publicado: (2025)
A Variational Approach for Mitigating Entity Bias in Relation Extraction
por: Mensah, Samuel, et al.
Publicado: (2025)
por: Mensah, Samuel, et al.
Publicado: (2025)
Grounding LLM Reasoning with Knowledge Graphs
por: Amayuelas, Alfonso, et al.
Publicado: (2025)
por: Amayuelas, Alfonso, et al.
Publicado: (2025)
Distill and Align Decomposition for Enhanced Claim Verification
por: Magomere, Jabez, et al.
Publicado: (2026)
por: Magomere, Jabez, et al.
Publicado: (2026)
The Influence of Biomedical Research on Future Business Funding: Analyzing Scientific Impact and Content in Industrial Investments
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
por: Khanmohammadi, Reza, et al.
Publicado: (2024)
Conservative Bias in Large Language Models: Measuring Relation Predictions
por: Aguda, Toyin, et al.
Publicado: (2025)
por: Aguda, Toyin, et al.
Publicado: (2025)
Detecting Non-Membership in LLM Training Data via Rank Correlations
por: Shetty, Pranav, et al.
Publicado: (2026)
por: Shetty, Pranav, et al.
Publicado: (2026)
Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency
por: Aguda, Toyin, et al.
Publicado: (2024)
por: Aguda, Toyin, et al.
Publicado: (2024)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
por: Shetty, Pranav, et al.
Publicado: (2025)
por: Shetty, Pranav, et al.
Publicado: (2025)
Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
por: Khanmohammadi, Reza, et al.
Publicado: (2026)
Calibrating LLM Confidence by Probing Perturbed Representation Stability
por: Khanmohammadi, Reza, et al.
Publicado: (2025)
por: Khanmohammadi, Reza, et al.
Publicado: (2025)
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
por: Fons, Elizabeth, et al.
Publicado: (2025)
por: Fons, Elizabeth, et al.
Publicado: (2025)
DocLLM: A layout-aware generative language model for multimodal document understanding
por: Wang, Dongsheng, et al.
Publicado: (2023)
por: Wang, Dongsheng, et al.
Publicado: (2023)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
por: Wu, Yunze, et al.
Publicado: (2025)
por: Wu, Yunze, et al.
Publicado: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
por: Liu, Chengzhi, et al.
Publicado: (2026)
por: Liu, Chengzhi, et al.
Publicado: (2026)
FinResearchBench: A Logic Tree based Agent-as-a-Judge Evaluation Framework for Financial Research Agents
por: Sun, Rui, et al.
Publicado: (2025)
por: Sun, Rui, et al.
Publicado: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
por: Haque, Mirazul, et al.
Publicado: (2025)
por: Haque, Mirazul, et al.
Publicado: (2025)
ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images
por: Sibue, Mathieu, et al.
Publicado: (2026)
por: Sibue, Mathieu, et al.
Publicado: (2026)
GenPlanX. Generation of Plans and Execution
por: Borrajo, Daniel, et al.
Publicado: (2025)
por: Borrajo, Daniel, et al.
Publicado: (2025)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
por: Zhu, Fengbin, et al.
Publicado: (2025)
por: Zhu, Fengbin, et al.
Publicado: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
por: Agrawal, Yogesh, et al.
Publicado: (2026)
por: Agrawal, Yogesh, et al.
Publicado: (2026)
FinSight: Towards Real-World Financial Deep Research
por: Jin, Jiajie, et al.
Publicado: (2025)
por: Jin, Jiajie, et al.
Publicado: (2025)
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
por: Liu, Shu, et al.
Publicado: (2024)
por: Liu, Shu, et al.
Publicado: (2024)
Bayesian Hierarchical Model Replication Study on Field Research Station Systems in Ghana,
por: Mensah, Yahaya, et al.
Publicado: (2001)
por: Mensah, Yahaya, et al.
Publicado: (2001)
PaperBench: Evaluating AI's Ability to Replicate AI Research
por: Starace, Giulio, et al.
Publicado: (2025)
por: Starace, Giulio, et al.
Publicado: (2025)
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
por: Zhu, Yiyun, et al.
Publicado: (2026)
por: Zhu, Yiyun, et al.
Publicado: (2026)
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
por: Shu, Dong, et al.
Publicado: (2025)
por: Shu, Dong, et al.
Publicado: (2025)
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
por: Hou, Yutao, et al.
Publicado: (2026)
por: Hou, Yutao, et al.
Publicado: (2026)
Does Financial Literacy Influence Investment Decisions Among Retail Investors? The Moderating Role of Behavioral Biases
por: Mensah Marfo, et al.
Publicado: (2026)
por: Mensah Marfo, et al.
Publicado: (2026)
Enhancing Investment Analysis: Optimizing AI-Agent Collaboration in Financial Research
por: Han, Xuewen, et al.
Publicado: (2024)
por: Han, Xuewen, et al.
Publicado: (2024)
FinReflectKG: Agentic Construction and Evaluation of Financial Knowledge Graphs
por: Arun, Abhinav, et al.
Publicado: (2025)
por: Arun, Abhinav, et al.
Publicado: (2025)
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
por: Dimino, Fabrizio, et al.
Publicado: (2025)
por: Dimino, Fabrizio, et al.
Publicado: (2025)
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
por: Lu, Jiaxuan, et al.
Publicado: (2026)
por: Lu, Jiaxuan, et al.
Publicado: (2026)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
por: Choi, Chanyeol, et al.
Publicado: (2025)
por: Choi, Chanyeol, et al.
Publicado: (2025)
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
por: Malarkkan, Arun Vignesh, et al.
Publicado: (2026)
por: Malarkkan, Arun Vignesh, et al.
Publicado: (2026)
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
por: Zmigrod, Ran, et al.
Publicado: (2024)
por: Zmigrod, Ran, et al.
Publicado: (2024)
EXP-Bench: Can AI Conduct AI Research Experiments?
por: Kon, Patrick Tser Jern, et al.
Publicado: (2025)
por: Kon, Patrick Tser Jern, et al.
Publicado: (2025)
Factors Affecting Nurses, Midwives and Allied Health Professionals' Ability to Engage With Research
por: Parveen Ali, et al.
Publicado: (2026)
por: Parveen Ali, et al.
Publicado: (2026)
Ejemplares similares
-
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
por: Singh, Kuldeep, et al.
Publicado: (2024) -
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
por: Magomere, Jabez, et al.
Publicado: (2025) -
A Variational Approach for Mitigating Entity Bias in Relation Extraction
por: Mensah, Samuel, et al.
Publicado: (2025) -
Grounding LLM Reasoning with Knowledge Graphs
por: Amayuelas, Alfonso, et al.
Publicado: (2025) -
Distill and Align Decomposition for Enhanced Claim Verification
por: Magomere, Jabez, et al.
Publicado: (2026)