Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zichen, Chen, Jiaao, Chen, Jianda, Sra, Misha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
por: Chen, Zichen, et al.
Publicado: (2023)
por: Chen, Zichen, et al.
Publicado: (2023)
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
por: Benhenda, Mostapha
Publicado: (2026)
por: Benhenda, Mostapha
Publicado: (2026)
Large Language Models in Finance: A Survey
por: Li, Yinheng, et al.
Publicado: (2023)
por: Li, Yinheng, et al.
Publicado: (2023)
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
por: Biancotti, Claudia, et al.
Publicado: (2024)
por: Biancotti, Claudia, et al.
Publicado: (2024)
Designing Heterogeneous LLM Agents for Financial Sentiment Analysis
por: Xing, Frank
Publicado: (2024)
por: Xing, Frank
Publicado: (2024)
Reasoning Models Ace the CFA Exams
por: Patel, Jaisal, et al.
Publicado: (2025)
por: Patel, Jaisal, et al.
Publicado: (2025)
Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks
por: Wang, Julian Junyan, et al.
Publicado: (2025)
por: Wang, Julian Junyan, et al.
Publicado: (2025)
The Agentic Regulator: Risks for AI in Finance and a Proposed Agent-based Framework for Governance
por: Kurshan, Eren, et al.
Publicado: (2025)
por: Kurshan, Eren, et al.
Publicado: (2025)
UniFinEval: Towards Unified Evaluation of Financial Multimodal Models across Text, Images and Videos
por: Yang, Zhi, et al.
Publicado: (2026)
por: Yang, Zhi, et al.
Publicado: (2026)
StakeBench: Evaluating Language Understanding Grounded in Market Commitment
por: Pei, Yunhua, et al.
Publicado: (2026)
por: Pei, Yunhua, et al.
Publicado: (2026)
Words That Unite The World: A Unified Framework for Deciphering Central Bank Communications Globally
por: Shah, Agam, et al.
Publicado: (2025)
por: Shah, Agam, et al.
Publicado: (2025)
A Financial Brain Scan of the LLM
por: Chen, Hui, et al.
Publicado: (2025)
por: Chen, Hui, et al.
Publicado: (2025)
Artificial Intelligence and Systemic Risk: A Unified Model of Performative Prediction, Algorithmic Herding, and Cognitive Dependency in Financial Markets
por: Meng, Shuchen, et al.
Publicado: (2026)
por: Meng, Shuchen, et al.
Publicado: (2026)
NumLLM: Numeric-Sensitive Large Language Model for Chinese Finance
por: Su, Huan-Yi, et al.
Publicado: (2024)
por: Su, Huan-Yi, et al.
Publicado: (2024)
Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining
por: Mao, Qian'ang, et al.
Publicado: (2025)
por: Mao, Qian'ang, et al.
Publicado: (2025)
AI Patents in the United States and China: Measurement, Organization, and Knowledge Flows
por: Fang, Hanming, et al.
Publicado: (2026)
por: Fang, Hanming, et al.
Publicado: (2026)
Workflow is All You Need: Escaping the "Statistical Smoothing Trap" via High-Entropy Information Foraging and Adversarial Pacing
por: Jiang, Zhongjie
Publicado: (2025)
por: Jiang, Zhongjie
Publicado: (2025)
Financial Statement Analysis with Large Language Models
por: Kim, Alex, et al.
Publicado: (2024)
por: Kim, Alex, et al.
Publicado: (2024)
Personalized Chain-of-Thought Summarization of Financial News for Investor Decision Support
por: Zhang, Tianyi, et al.
Publicado: (2025)
por: Zhang, Tianyi, et al.
Publicado: (2025)
When Valid Signals Fail: Regime Boundaries Between LLM Features and RL Trading Policies
por: Yang, Zhengzhe
Publicado: (2026)
por: Yang, Zhengzhe
Publicado: (2026)
Engaging with AI: How Interface Design Shapes Human-AI Collaboration in High-Stakes Decision-Making
por: Chen, Zichen, et al.
Publicado: (2025)
por: Chen, Zichen, et al.
Publicado: (2025)
A Scoping Review of ChatGPT Research in Accounting and Finance
por: Dong, Mengming Michael, et al.
Publicado: (2024)
por: Dong, Mengming Michael, et al.
Publicado: (2024)
Explaining the Unexplainable: A Systematic Review of Explainable AI in Finance
por: Mohsin, Md Talha, et al.
Publicado: (2025)
por: Mohsin, Md Talha, et al.
Publicado: (2025)
NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance
por: Lee, Hanwool, et al.
Publicado: (2025)
por: Lee, Hanwool, et al.
Publicado: (2025)
A Survey of Large Language Models in Finance (FinLLMs)
por: Lee, Jean, et al.
Publicado: (2024)
por: Lee, Jean, et al.
Publicado: (2024)
FinRobot: Generative Business Process AI Agents for Enterprise Resource Planning in Finance
por: Yang, Hongyang, et al.
Publicado: (2025)
por: Yang, Hongyang, et al.
Publicado: (2025)
Harnessing Earnings Reports for Stock Predictions: A QLoRA-Enhanced LLM Approach
por: Ni, Haowei, et al.
Publicado: (2024)
por: Ni, Haowei, et al.
Publicado: (2024)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
por: Wu, Qiming, et al.
Publicado: (2024)
por: Wu, Qiming, et al.
Publicado: (2024)
Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions
por: Lin, Ryan Y., et al.
Publicado: (2024)
por: Lin, Ryan Y., et al.
Publicado: (2024)
Dissecting AI Trading: Behavioral Finance and Market Bubbles
por: Ouyang, Shumiao, et al.
Publicado: (2026)
por: Ouyang, Shumiao, et al.
Publicado: (2026)
Mixing It Up: The Cocktail Effect of Multi-Task Fine-Tuning on LLM Performance -- A Case Study in Finance
por: Brief, Meni, et al.
Publicado: (2024)
por: Brief, Meni, et al.
Publicado: (2024)
A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges
por: Nie, Yuqi, et al.
Publicado: (2024)
por: Nie, Yuqi, et al.
Publicado: (2024)
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
por: Ye, Yuxuan, et al.
Publicado: (2026)
por: Ye, Yuxuan, et al.
Publicado: (2026)
Do LLM Personas Dream of Bull Markets? Comparing Human and AI Investment Strategies Through the Lens of the Five-Factor Model
por: Borman, Harris, et al.
Publicado: (2024)
por: Borman, Harris, et al.
Publicado: (2024)
Finance Language Model Evaluation (FLaME)
por: Matlin, Glenn, et al.
Publicado: (2025)
por: Matlin, Glenn, et al.
Publicado: (2025)
LMExplainer: Grounding Knowledge and Explaining Language Models
por: Chen, Zichen, et al.
Publicado: (2023)
por: Chen, Zichen, et al.
Publicado: (2023)
LLM Agents for Combinatorial Efficient Frontiers: Investment Portfolio Optimization
por: Paquette-Greenbaum, Simon, et al.
Publicado: (2026)
por: Paquette-Greenbaum, Simon, et al.
Publicado: (2026)
AI-Driven Alpha Decay: Algorithmic Homogenization, Reflexive Signal Erosion, and the Paradox of Intelligent Markets
por: Meng, Shuchen, et al.
Publicado: (2026)
por: Meng, Shuchen, et al.
Publicado: (2026)
The LLM Pro Finance Suite: Multilingual Large Language Models for Financial Applications
por: Caillaut, Gaëtan, et al.
Publicado: (2025)
por: Caillaut, Gaëtan, et al.
Publicado: (2025)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
por: Wang, Yan, et al.
Publicado: (2025)
por: Wang, Yan, et al.
Publicado: (2025)
Ejemplares similares
-
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
por: Chen, Zichen, et al.
Publicado: (2023) -
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
por: Benhenda, Mostapha
Publicado: (2026) -
Large Language Models in Finance: A Survey
por: Li, Yinheng, et al.
Publicado: (2023) -
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
por: Biancotti, Claudia, et al.
Publicado: (2024) -
Designing Heterogeneous LLM Agents for Financial Sentiment Analysis
por: Xing, Frank
Publicado: (2024)